About This Role
We are seeking a Senior Data Platform Engineer to build and operate the core data infrastructure making CertifyOS the definitive source of truth for provider data. You will own pipeline reliability, performance optimization, and collaborate closely with engineering, product, and ML teams.
Key Responsibilities:Create and maintain distributed real-time (Kafka, Kinesis) and batch pipelines using Python and PySpark.Maintain schema integrity and storage strategy across OLTP (Aurora PostgreSQL), OLAP (Redshift), and NoSQL (MongoDB) systems.Ensure data quality, contractual compliance, and observability for all built solutions.Tune queries to optimize throughput, latency, and operational costs.Lead entity resolution efforts within our MDM system for provider data accuracy.Establish robust data foundations supporting AI/ML workloads.Requirements:5+ years of experience building production-grade data systems in regulated environments (Healthcare/PHI).Expert-level proficiency in Python and PySpark with deep knowledge of distributed system fundamentals (consistency, partitioning, idempotency).Proven track record of shipping performance improvements to pipelines.Solid experience with event-driven architecture and data observability practices.Familiarity with AI-assisted development workflows using tools like Cursor or Claude Code.Nice-to-Have:Experience with entity resolution, CDC patterns, or serving ML workloads.Knowledge of orchestration tools such as dbt, Airflow, or AWS Glue.Tech Stack:Pipeline: Python/PySpark | Streaming: Kafka/Kinesis/SNS-SQS | Storage: Aurora PostgreSQL, Redshift, MongoDB Atlas | Infrastructure: AWS (EKS), Docker