About Us
STACK is a leading provider of cloud-based construction estimating and takeoff software solutions, committed to helping businesses transform through innovative solutions. We pride ourselves on fostering a collaborative, dynamic environment where team members have the opportunity to grow and make a real impact.
About the Role
We are hiring a Senior Computer Vision Engineer to improve our computer vision system for construction documents. This role focuses on detection and segmentation quality, geometric accuracy, and production-grade pipelines that work reliably with messy real-world PDFs and drawings.
What You’ll Do
- Design, build, and improve end-to-end detection and segmentation pipelines for document images.
- Improve document ingestion for PDFs and other unstructured files, including parsing, rendering, and handling of multi-page and multi-view content.
- Increase model accuracy, boundary and geometric quality, and overall reliability of predictions for downstream use.
- Integrate modern detection and segmentation models into production workflows and build the post-processing that turns model output into structured, usable geometry.
- Define evaluation metrics, investigate failure cases, and drive continuous quality improvements.
- Optimize latency, reliability, and cost across the inference and post-processing stack.
- Own training infrastructure, dataset curation, annotation quality, and continuous-improvement loops.
- Make architectural decisions and own system quality end to end.
What You Bring
- 5+ years experience building computer vision systems: detection, segmentation, or structured geometry extraction, used high volume in production.
- Experience working with messy, real-world image data or large unstructured visual datasets.
- Strong understanding of detection and segmentation tradeoffs, including model architecture choices, training data design, and post-processing.
- Ability to measure system performance with evaluation, testing, and production metrics.
- Ability to explain failure modes clearly and improve systems through debugging, dataset work, and iteration.
- Experience with multimodal models (vision-language models, document AI systems)
- Understanding of grounding — linking model outputs to source data or coordinates
- Backend engineering experience, including APIs, async processing, and scalable GPU services.
Additional Preferred Qualifications
- Experience with polygon or mask post-processing, geometric regularization, or CAD-style structured output.
- Experience with layout-aware document processing, PDF vector extraction, or combining raster and vector signals.
- Background in document-heavy CV domains such as construction, real estate, medical imaging, geospatial, or similar workflows.
- Experience optimizing inference cost and latency at scale.
- Familiarity with open-source detection / segmentation ecosystems, training infrastructure, or model serving.
What Success Looks Like
- Accurate, geometrically correct predictions suitable for downstream measurement, and analysis use.
- Fast, reliable inference across large and messy real-world document sets.
- Clear quality metrics and a repeatable improvement loop.
- Systems that perform consistently under real-world production constraints.
What This Role is Not
- Not a model-training-only role. You'll own data, training, post-processing, and serving.
- Not a research-only role.
- Not a plug-and-play CV tools environment.
Why Join STACK?
- Opportunity to work in a fast-paced, growth-oriented, remote-first environment.
- Async-friendly environment with a focus on ownership and deep work.
- Be part of a dynamic, supportive team where your contributions are valued.
At STACK, our values shape how we work, collaborate, and serve our customers:
- Radical Honesty:Communicate directly, respectfully, and transparently—even when conversations are difficult. Give and receive fe