We are looking for Data Engineer candidates for a project delivered through micro1.
What you'll do
- Design, build, and maintain scalable data pipelines to ingest, process, and transform large-scale datasets from multiple sources.
- Develop and optimize distributed data processing workflows using Spark and cloud-native technologies.
- Build and maintain data storage solutions across SQL and NoSQL systems, ensuring scalability, performance, and reliability.
- Design and implement data architectures on AWS to support high-volume data ingestion, processing, and distribution.
- Write efficient Python and SQL code to extract, transform, validate, and analyze large datasets.
- Ensure data quality, integrity, monitoring, and operational reliability across data pipelines and storage layers.
- Collaborate with AI researchers, data scientists, and engineering teams to support data-intensive applications and experimentation.
- Implement automation, orchestration, and monitoring workflows to support scalable and efficient data operations.
What you need
- Strong proficiency in Python, SQL, and distributed data processing frameworks such as Apache Spark.
- Hands-on experience with AWS data services and cloud-native data architectures.
- Experience working with both SQL and NoSQL databases.
- Experience managing and processing large-scale datasets in distributed environments.
- Strong understanding of data partitioning, performance optimization, and scalable data architectures
Nice to have
- Exposure to AI/ML workflows or research environments.
- Experience with data visualization tools such as Matplotlib, Seaborn, or Plotly.
- Familiarity with LLM-related data workflows (datasets for training, evaluation, or prompt experimentation).
Expertise
Who you work with
Project and contracting process: micro1. Applications continue on the provider's website.

