at Cherokee Federal
ATA is seeking a Data Scientist to support data pipeline development, validation, and analysis within a cloud-based Health IT data platform. This role is hands-on and delivery-focused, with an emphasis on building reliable, reproducible data workflows using SQL, Python, and PySpark in an Azure Synapse environment.
A core expectation of this role is the ability to work across the full data lifecycle, from ingestion through transformation to final dataset delivery, while maintaining data quality and traceability. The ideal candidate is comfortable debugging data issues end-to-end, understands how data structure and join logic impact outputs, and applies disciplined validation and documentation practices. This role also supports exploratory data analysis and the development of derived datasets to enable analytics and downstream use cases. The position will work extensively with healthcare data originating from EHR systems and interface feeds, including HL7 v2 and FHIR data, clinical terminology, and source-to-target data mappings.
Key Responsibilities
Data Pipeline Development
PySpark.
ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | www.ata-llc.com
Advanced Technology Applications
debugging and documentation.
JSON/NDJSON, and Parquet.
feeds while preserving source lineage and clinical context
Data Quality and Validation
detection, schema validation, and allowed value enforcement.
transformation errors across pipeline stages.
fixes, and ensuring data reliability prior to downstream use.
systems during data conversion and migration activities.
Data Analysis and Dataset Development
concerns.
cases.
metrics.
Data Modeling and Structure Awareness
and composite keys, and table relationships.
many) impact row counts and outputs.
Debugging and Troubleshooting
clear and consistent manner.
integrity, and alignment with requirements.
ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | www.ata-llc.com
Advanced Technology Applications
members.
Minimum Qualifications
-