Join Xperteez Technology as a Remote Data Engineer to build and scale the data infrastructure powering AI-driven products and research initiatives. In this role you’ll design, build, and maintain distributed pipelines, manage large-scale datasets across cloud environments, and craft reliable data architectures on AWS using Python, SQL, Spark, and a mix of SQL and NoSQL stores. You’ll prioritize data quality, monitoring, automation, and collaborate with AI researchers and data scientists to support experimentation and model development at scale. To apply, highlight hands-on experience with ETL, orchestration, and performance optimization, quantify impact with throughput, latency, and cost savings, and share links to projects or GitHub.
Job Title: Data Engineer
Job Type: Full-time
Location: Remote
We are looking for a Data Engineer to build and scale the data infrastructure that powers AI-driven products and research initiatives. In this role, you will develop distributed data pipelines, manage large-scale datasets across cloud environments, and design reliable data systems that support data processing, experimentation, and model development at scale.
Required Skills
Python
SQL
AI/ML
Spark
AWS
Key Responsibilities
Design, build, and maintain scalable data pipelines to ingest, process, and transform large-scale datasets from multiple sources.
Develop and optimize distributed data processing workflows using Spark and cloud-native technologies.
Build and maintain data storage solutions across SQL and NoSQL systems, ensuring scalability, performance, and reliability.
Design and implement data architectures on AWS to support high-volume data ingestion, processing, and distribution.
Write efficient Python and SQL code to extract, transform, validate, and analyze large datasets.
Ensure data quality, integrity, monitoring, and operational reliability across data pipelines and storage layers.
Collaborate with AI researchers, data scientists, and engineering teams to support data-intensive applications and experimentation.
Implement automation, orchestration, and monitoring workflows to support scalable and efficient data operations.
Required Skills and Qualifications
Strong proficiency in Python, SQL, and distributed data processing frameworks such as Apache Spark.
Hands-on experience with AWS data services and cloud-native data architectures.
Experience working with both SQL and NoSQL databases.
Experience managing and processing large-scale datasets in distributed environments.
Strong understanding of data partitioning, performance optimization, and scalable data architectures
Nice to Have
Exposure to AI/ML workflows or research environments.
Experience with data visualization tools such as Matplotlib, Seaborn, or Plotly.
Familiarity with LLM-related data workflows (datasets for training, evaluation, or prompt experimentation).
Compensation & Benefits Notice
All employees are eligible for equity compensation, and employees may also receive performance-based bonuses, dependent on role and subject to company policies. Comprehensive benefits package, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, remote-first workforce.
Disclaimer
The information contained in this job posting, including but not limited to role responsibilities, qualifications, compensation, and benefits, is provided for informational purposes only and does not constitute a binding offer of employment. micro1 reserves the right to amend, modify, or withdraw any portion of this posting at its sole discretion and without prior notice. All employment decisions are made in accordance with applicable laws and regulations.