The OpportunityWe’re looking for a Staff Data Platform Engineer to architect and scale the data backbone powering next-generation AI models in robotics and real-world environments.This role sits at the intersection of distributed systems, multimodal data processing, and applied machine learning, with a strong focus on building high-quality datasets for robotic foundation models. You will ensure that data pipelines, infrastructure, and data strategy directly translate into measurable improvements in model performance.Your ResponsibilitiesDrive the model–data loop by connecting application requirements with data collection, and translating model failures into data-driven improvements through collection, curation, and augmentationBuild and scale distributed data pipelines (Ray/Anyscale or similar) for TB-scale video, sensor, and robotics datasetsDesign multimodal data schemas aligning video, actions, and high-frequency sensor streamsDevelop Python tooling for data quality, including cleaning, anomaly detection, and dataset versioningOwn dataset quality and coverage, including annotation workflows, data diversity, and storage trade-offsLead a small team and coordinate with data providers and annotation vendorsOversee real-world data collection, including technical setup, compliance, and secure data handlingTechnologiesPython (advanced, production-grade)Ray / Anyscale or Apache SparkAWS / GCP for large-scale data and GPU training pipelinesVideo and sensor data formats (H.264/H.265, ROS bags, MCAP)PyTorch, NumPyDVC, LakeFS or similar data versioning toolsDistributed data processing and storage systemsMust Have10+ years in Data/ML Engineering, including 6+ years in a senior or lead roleExperience with large-scale real-world data (robotics, autonomous systems, or video AI)Strong experience with Ray/Anyscale or Spark for distributed pipelinesAdvanced Python (performance, concurrency, ML stack like NumPy/PyTorch)Experience working with video and sensor data formats (e.g., H.264/H.265, ROS bags, MCAP)Experience building scalable data pipelines for GPU-based training workloads (AWS/GCP)Experience with data versioning tools such as DVC or LakeFSProven experience owning systems and mentoring engineersNice to HaveExperience building datasets for multimodal foundation models (VLA, VLM or similar)Robotics fundamentals (sensor synchronization, 3D transforms)Experience with active learning or data-centric ML workflowsCompetitive compensation packageVarious employee subsidies and perks, including public transportation and WellpassWork with a world-class team in a flat hierarchy, with direct collaboration alongside the founders and engineering teamOpportunity to make a real impact by working on cutting-edge robotics and AI systemsFast growth potential in a rapidly evolving company and industryInternational office environment with English as the official working languageRecruiting ProcessYour recruiting partner for this role is Madhulika (she/her). The process includes a screening interview, two virtual technical interviews, and an onsite visit to our office in Munich to meet with the team.We hire across backgrounds, identities, and experiences, and we are committed to a workplace where everyone belongs. Discrimination has no place here.If you need any accommodations during the recruiting process, just reach out to your recruiting partner. #J-18808-Ljbffr
Suche nach anderen Stellen
Stellenbeschreibung
Jobalert für diese Suche erstellen
Staff Data Platform Engineer (F/M/D) • München, bayern, Germany