Atyeti Inc•2h ago
LinkedIn
Lead BigData Engineer with
India
Senior Level
Full Job Description
Role: Lead BigData Engineer
About Us
We are seeking a seasoned engineer to lead our Hadoop-to-Databricks modernization initiative. You will drive the migration of large-scale data platforms from legacy on-premise solutions to cloud-native architectures, ensuring seamless performance and scalability.
Key Responsibilities
- Migration Strategy & Execution: Design and lead comprehensive strategies for migrating HDFS, Hive, Spark, MapReduce workloads to Databricks. Perform source-to-target mapping and data profiling.
- Databricks Optimization: Expertly refactor existing Spark codebases (Scala/Python) for the Databricks environment. Optimize job performance by tuning partitioning strategies, managing memory utilization, handling shuffles efficiently, and leveraging Delta Lake features like ACID transactions, schema evolution, and time travel.
- Data Pipeline Development: Build scalable ETL/ELT pipelines using Spark SQL, PySpark, and Databricks Workflows. Implement robust job orchestration for complex data dependencies.
- Cloud Storage Integration: Migrate terabytes of historical data from HDFS to cloud storage solutions like Azure Data Lake Storage (ADLS), AWS S3, or Google Cloud Storage.
- Data Quality & Governance: Implement validation frameworks, reconciliation processes, and error-handling mechanisms. Ensure all migrated workloads adhere to enterprise security, governance, audit, and compliance standards essential for banking environments.
- Modernization of Legacy SQL: Convert legacy Hive/Spark SQL logic into optimized Databricks SQL queries and Delta Live Tables where appropriate.
- Cross-Functional Collaboration: Partner closely with data architects, scientists, business analysts, and application teams to understand requirements and deliver high-value solutions.
- DevOps & CI/CD: Implement continuous integration and deployment pipelines for Databricks notebooks and jobs using Git-based workflows. Automate testing environments for migration validation.
- Troubleshooting: Actively monitor cluster health, troubleshoot performance bottlenecks in Spark clusters and data pipelines to ensure 99.9% availability.
Company
Atyeti Inc
Atyeti is a leading professional services and technology consulting firm headquartered in Princeton, New Jersey, with global offices spanning USA, UK, Ireland, Switzerland, Poland, India, Singapore, M...
India
Posted on LinkedIn