HO - Data Engineer Specialist
Data Engineer Specialist is responsible for designing, implementing, managing, and optimizing enterprise Data Lake environments that support large-scale data storage, processing, analytics, artificial intelligence (AI), and machine learning (ML) initiatives. This role ensures that data is efficiently ingested, governed, secured, and made available for business intelligence, analytics, and advanced data science use cases.
Collaborates with Data Engineers, Data Scientists, Analytics teams, Cloud Architects, and business stakeholders to create scalable, high-performance data platforms that enable data-driven decision-making across the organization.
Key Responsibilities
Data Lake Architecture & Design
- Design and develop enterprise Data Lake and Lakehouse architectures.
- Define data storage structures for raw, curated, and consumption layers.
- Establish data partitioning, indexing, and lifecycle management strategies.
- Ensure scalability, reliability, and performance of data storage environments.
- Participate in data platform modernization and cloud transformation initiatives.
Data Ingestion & Integration
- Build and maintain data ingestion pipelines from various sources, including:
- ERP systems
- CRM platforms
- IoT devices
- Databases
- APIs
- Third-party applications
- Support both batch and real-time data ingestion processes.
- Implement change data capture (CDC) and streaming data solutions.
- Ensure seamless integration with enterprise ecosystems.
Data Lake Operations
- Manage and monitor data lake platforms and storage environments.
- Automate operational tasks and platform maintenance activities.
- Monitor data load performance, storage utilization, and processing efficiency.
- Troubleshoot and resolve platform-related issues.
Data Governance & Quality
- Implement data governance frameworks and best practices.
- Manage metadata, cataloging, lineage, and business glossary initiatives.
- Develop data quality monitoring and validation processes.
- Ensure compliance with organizational data standards and policies.
- Support Master Data Management (MDM) and data stewardship activities.
Security & Compliance
- Implement role-based access controls (RBAC) and data security policies.
- Ensure compliance with data privacy regulations and corporate governance standards.
- Monitor security vulnerabilities and implement remediation measures.
- Support security audits and risk assessments.
Cloud Data Platform Management
- Manage cloud-native data lake solutions.
- Optimize storage and compute resource utilization.
- Implement infrastructure automation and Infrastructure as Code (IaC).
- Support hybrid and multi-cloud data architectures.
Advanced Analytics Enablement
- Prepare and optimize datasets for analytics, reporting, and AI/ML projects.
- Enable self-service analytics through governed data access.
- Collaborate with Data Scientists and BI developers to improve data accessibility.
- Support implementation of Lakehouse and modern analytics architectures.
Documentation & Continuous Improvement
- Create and maintain architecture diagrams, standards, and operational procedures.
- Evaluate emerging technologies and recommend enhancements.
- Promote automation, operational excellence, and engineering best practices.
- Provide technical leadership and mentoring to junior team members.
Required Qualifications
Education
- Bachelor‘s degree in Computer Science, Information Technology, Data Engineering, Data Science, Software Engineering, or a related field.
Experience
- At least 5 years of experience in Data Engineering, Data Platform Administration, Data Warehousing, or Big Data solutions.
- 2+ years of hands-on experience with Data Lake or Lakehouse platforms.
- Experience managing cloud-based analytics and storage environments.
- Proven experience designing scalable enterprise data architectures.
Technical Skills
Data Lake & Lakehouse Platforms
- Azure Data Lake Storage (ADLS)
- Azure Microsoft Fabric OneLake
- Amazon S3 Data Lake
- Google Cloud Storage
- Databricks Lakehouse Platform
- Delta Lake
- Apache Iceberg
- Apache Hudi
Data Engineering
- ETL/ELT Frameworks
- Apache Spark
- Apache Kafka
- Azure Data Factory (ADF)
- Apache Airflow
- Databricks Workflows
- Data Pipeline Orchestration
Cloud Platforms
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
Databases & Warehouses
- Snowflake
- Azure Synapse Analytics
- Amazon Redshift
- Google BigQuery
- Oracle
- SQL Server
- PostgreSQL
Programming Languages
- SQL
- Python
- Scala
- PySpark
- Shell Scripting
Data Governance & Security
- Data Catalog Solutions
- Microsoft Purview
- Collibra
- Metadata Management
- Data Lineage
- Data Classification
- Access Control & Encryption
DevOps & Automation
- Git
- GitHub / GitLab
- CI/CD Pipelines
- Docker
- Kubernetes
- Terraform
- Infrastructure as Code (IaC)
Soft Skills
- Strong analytical and problem-solving abilities.
- Excellent communication and stakeholder management skills.
- Ability to work cross-functionally with technical and business teams.
- Strong organizational and documentation capabilities.
- Proactive mindset focused on automation and continuous improvement.
- Ability to manage multiple priorities in fast-paced environments.