AI Intern – Vision-Language-Action (VLA) & Data
RIVRJobsCloseBy Editorial Insights
RIVR, an Amazon robotics company and ETH Zurich spin-off, seeks an AI Intern for Vision-Language-Action and Data to help build data pipelines and train multi-modal models for real world robots. This onsite Zurich role sits at the intersection of data engineering and model training, with hands on work on VLA systems, data visualization tools, and debugging model performance. You will curate large multi-modal datasets, work with senior engineers to implement data strategies, and test algorithms in simulation and on hardware. To apply, emphasize Python and PyTorch, computer vision DL experience, data tooling (NumPy, Pandas, OpenCV), and relevant projects; highlight collaboration, problem solving, and eagerness to learn. Important: Schengen citizenship required, unless you are an ETHZ/EPFL student intern; in person presence required.
RIVR, part of Amazon is a robotics company pioneering Physical AI through real-world doorstep delivery. Founded in 2024 as an ETH Zurich spin-off, RIVR, part of Amazon developed wheeled-legged robots designed to operate in complex, unstructured environments such as stairs, gates, doors, and uneven urban terrain. We believe that achieving general physical intelligence requires solving real customer problems in the real world, where robots can learn from rich operational data at scale.
Following our acquisition by Amazon in March 2026, we are continuing this mission with greater reach and speed. By combining custom robot hardware, onboard autonomy, and cloud-based coordination, RIVR, part of Amazon is building the next generation of safe, reliable autonomous robots for last-mile delivery.
Important Notice: For this position, we can unfortunately only accept applications from citizens of Schengen Area countries. This restriction does not apply to ETHZ and EPFL students who are required to complete compulsory internships as part of their studies.Job Description
As an AI Intern - VLA & Data, you will assist the AI engineering team in developing data pipelines and training Vision-Language-Action (VLA) models for robotic systems. In this role, you will sit at the intersection of data engineering and model training, assisting the team in solving the "data bottleneck" in embodied AI. Additionally, your responsibilities will include building tools to analyze datasets, visualize model predictions, and debug performance. You will work closely with senior engineers to understand, curate, and visualize the massive amounts of multi-modal data our fleet generates. This role offers hands-on experience within a dynamic and collaborative environment. We are committed to finding and nurturing exceptional talent; our internships are a key pathway to recruiting outstanding graduates who can make a significant impact in our team.
What you’ll be doing
-
Support the team in processing, curating, and analyzing multi-modal sensor data for training VLA models.
-
Get hands-on experience with state-of-the-art VLA models.
-
Assist in developing software tools to visualize data and debug model performance.
-
Work closely with senior engineers to implement data strategies that improve the robustness of our robotic systems.
-
Help integrate software components to evaluate algorithms in simulation and on hardware.
-
Engage in continuous learning and gain exposure to recent developments in VLA, self-supervised learning, and generative AI.
What you must have
-
At least BSc in Computer Science, Robotics, Machine Learning, or a related field.
-
Proficiency in Python and experience with deep learning frameworks (preferably PyTorch).
-
Hands-on experience (either through coursework, previous internships, or other projects) with deep learning for computer vision.
-
Familiarity with data manipulation libraries (e.g., NumPy, Pandas, OpenCV).
-
Strong problem-solving skills and an eagerness to work with complex, real-world data.
-
Eagerness to learn and contribute in a collaborative team environment.
Get some bonus points
-
An MSc or ongoing PhD in a related field.
-
Previous experience working with large-scale image or video datasets.
-
Knowledge or experience with transformers, Vision-Language Models and/or VLA’s.
-
Experience with 3D geometry, camera projections, or sensor fusion.
-
Experience working with robotic systems
-
Experience in motion/action prediction in robotics context.