Data Engineer
Xperteez TechnologyXperteez Technology is looking for a Data Engineer to build and scale the data infrastructure that powers AI-driven products and research initiatives. The role sits in International City, AE with a hybrid work arrangement, blending in-person collaboration with remote work. The company is expanding its data capabilities to support experiments, model development, and data processing at scale.
The person in this role will shape the data backbone by designing distributed pipelines, managing large datasets across cloud environments, and creating reliable systems that enable data processing and experimentation. They’ll partner with AI researchers and data scientists to ensure data flows support research and product workflows, while keeping quality and reliability front and center.
If you enjoy turning data into scalable, well-governed platforms and working with cross-functional teams to advance AI initiatives, this opportunity could be a strong fit.
Making the data backbone for AI work
On a typical day, you’ll design, build, and maintain pipelines that ingest, process, and transform data from multiple sources, ensuring downstream teams can access clean, usable data for analysis and experimentation. You’ll also work on distributed data processing with Spark and cloud-native tools, shaping storage across SQL and NoSQL databases and crafting AWS-based architectures that handle high-volume data movement. Writing efficient Python and SQL code to extract, transform, validate, and analyse large datasets will be a core part of the role, as will monitoring data quality and the operational health of pipelines and storage layers. Collaboration with AI researchers, data scientists, and engineering peers will be frequent, and you’ll implement automation, orchestration, and monitoring to keep data operations scalable and reliable.
What you’ll bring to the role
- Proficiency in Python and SQL, with hands-on use of a distributed processing framework such as Spark
- Experience building and operating data solutions on AWS and with cloud-native data architectures
- Familiarity with both SQL and NoSQL databases in real-world projects
- History of handling large-scale datasets in distributed environments
- Strong grasp of data partitioning, performance tuning, and scalable architectural design
Nice to have, if you can add it
- Exposure to AI or ML workflows within research or product settings
- Experience using data visualization tools like Matplotlib, Seaborn, or Plotly
- Familiarity with data workflows related to large language models, including datasets for training, evaluation, or prompt experimentation
Compensation, benefits and the practical bits
The base salary for this full-time role ranges from 100,000.00 USD to 150,000.00 USD per year. In addition to the salary, employees may receive equity compensation and performance-based bonuses, subject to company policy and role requirements. The benefits package is comprehensive, including up to 100% reimbursement for health-insurance premiums, paid time off, a 401(K) plan with a company match, and additional benefits designed to support a high-performing, hybrid workforce.
As part of its approach to hiring, the company notes that AI tools are used to assist in screening and assessment, designed to complement human decision-making rather than replace it.
Tips for applicants
Put the must-have skills up front on your resume, especially Python, SQL, Spark, and AWS experience, and call out any work with SQL and NoSQL databases and large-scale datasets.
Back up the must-have claims with concrete examples of pipelines you’ve built, architectures you’ve deployed on AWS, and the data quality or reliability improvements you achieved through monitoring or automation.
Be prepared to walk through how you’ve collaborated with AI researchers or data scientists to support experiments or model development, including how data pipelines supported those efforts.
Ask a targeted question about the current AWS data stack, data governance practices, or how the team partners with AI researchers to run experiments and iterate quickly on data models.