AWS DATA ENGINEER & ETL EXPERT
Posted on:
4 hours ago
Vacancies:
1 Vacancy
Job Summary
We are looking for part time Freelancer for AWS Data Engineer Position.
Releavant Exp : 9
Mandatory Tech Stack required:
Source: Salesforce (ECRM & OSC)
Backup: Grax
Cloud: AWS (EC2 S3 CloudWatch DynamoDB RDS Secrets Manager ALB ASG)
Infrastructure as Code: Terraform
File Format: Parquet
Target: Enterprise Data Lake (EDL)
Schema: Blue Schema
Database: Blue Database
Query Language: SQL
JD:
1. Build and maintain AWS data pipelines
* Develop ETL/ELT pipelines using AWS Glue PySpark Python and SQL.
* Ingest data from sources like Salesforce databases APIs and S3.
* Load curated data into Amazon Redshift and data lake storage.
2. Optimize Athena and data lake performance
* Convert JSON/CSV data into Parquet.
* Use Snappy compression.
* Design proper partitioning strategies.
* Resolve split limit and performance issues.
* Optimize Athena query costs.
3. Manage modern data lake architecture
* Work with Apache Iceberg tables.
* Perform migrations from traditional Parquet tables.
* Support schema evolution time travel and ACID transactions.
4. Production support and troubleshooting
* Investigate Glue jobs that suddenly become slow.
* Debug Lambda timeouts.
* Fix missing records and data quality issues.
* Resolve Redshift performance problems.
* Perform root cause analysis (RCA).
5. Infrastructure as Code
* Build AWS infrastructure using Terraform.
* Create reusable modules.
* Manage Auto Scaling Groups ALBs IAM Lambda Secrets Manager S3 Redshift and DynamoDB.
* Troubleshoot Terraform state and production deployment issues.
6. Security
* Manage AWS Secrets Manager.
* Configure Lambda-based secret rotation.
* Ensure Terraform does not overwrite rotated passwords.
* Implement IAM least-privilege access.
7. Backup and Disaster Recovery
* Work with enterprise backup tools like Rubrik (their environment may use Grax for Salesforce).
* Validate backup jobs.
* Perform restores.
* Support disaster recovery testing.
* Verify restored data.
8. Data Quality
* Validate source and target record counts.
* Maintain audit/control tables (for example in DynamoDB).
* Check checksums and duplicate records.
* Troubleshoot data discrepancies reported by business users.
9. AWS Services
* S3
* Glue
* Athena
* Lambda
* Redshift
* DynamoDB
* EventBridge
* Step Functions
* CloudWatch
* Secrets Manager
* IAM
* Auto Scaling Groups
* Application Load Balancer
Releavant Exp : 9
Mandatory Tech Stack required:
Source: Salesforce (ECRM & OSC)
Backup: Grax
Cloud: AWS (EC2 S3 CloudWatch DynamoDB RDS Secrets Manager ALB ASG)
Infrastructure as Code: Terraform
File Format: Parquet
Target: Enterprise Data Lake (EDL)
Schema: Blue Schema
Database: Blue Database
Query Language: SQL
JD:
1. Build and maintain AWS data pipelines
* Develop ETL/ELT pipelines using AWS Glue PySpark Python and SQL.
* Ingest data from sources like Salesforce databases APIs and S3.
* Load curated data into Amazon Redshift and data lake storage.
2. Optimize Athena and data lake performance
* Convert JSON/CSV data into Parquet.
* Use Snappy compression.
* Design proper partitioning strategies.
* Resolve split limit and performance issues.
* Optimize Athena query costs.
3. Manage modern data lake architecture
* Work with Apache Iceberg tables.
* Perform migrations from traditional Parquet tables.
* Support schema evolution time travel and ACID transactions.
4. Production support and troubleshooting
* Investigate Glue jobs that suddenly become slow.
* Debug Lambda timeouts.
* Fix missing records and data quality issues.
* Resolve Redshift performance problems.
* Perform root cause analysis (RCA).
5. Infrastructure as Code
* Build AWS infrastructure using Terraform.
* Create reusable modules.
* Manage Auto Scaling Groups ALBs IAM Lambda Secrets Manager S3 Redshift and DynamoDB.
* Troubleshoot Terraform state and production deployment issues.
6. Security
* Manage AWS Secrets Manager.
* Configure Lambda-based secret rotation.
* Ensure Terraform does not overwrite rotated passwords.
* Implement IAM least-privilege access.
7. Backup and Disaster Recovery
* Work with enterprise backup tools like Rubrik (their environment may use Grax for Salesforce).
* Validate backup jobs.
* Perform restores.
* Support disaster recovery testing.
* Verify restored data.
8. Data Quality
* Validate source and target record counts.
* Maintain audit/control tables (for example in DynamoDB).
* Check checksums and duplicate records.
* Troubleshoot data discrepancies reported by business users.
9. AWS Services
* S3
* Glue
* Athena
* Lambda
* Redshift
* DynamoDB
* EventBridge
* Step Functions
* CloudWatch
* Secrets Manager
* IAM
* Auto Scaling Groups
* Application Load Balancer