Member of Technical Staff - World Models
Veeda AI · Hybrid
This listing is no longer verified as available.
MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.
Source: the employer's Ashby job board. Open the original listing for current details. Availability is not verified for this retained page.
Job details
- Work model
- Hybrid
- Salary
- Not listed by source
- Location
- Toronto
What the role asks for
What you'd do
- Generative World Model Architecture:
- Design, train, and scale action-conditioned video and latent dynamics models: diffusion transformers with rectified flow, block-causal autoregressive hybrids, causal video tokenizers whose compression ratio sets how far a rollout survives.
- Action Conditioning & Latent Actions:
- Condition rollouts on robot action chunks and camera trajectories, and recover pseudo-actions from unlabeled video with inverse dynamics and latent action models, so training is not capped by what teleoperation produced.
- Long-Horizon Stability & Memory:
- Attack compounding drift by training on the model's own rollouts (diffusion forcing, self-forcing) and carrying persistent scene state as 3D Gaussians or point maps, so minute-long rollouts stay coherent.
- Post-Training & Distillation:
- Fine-tune against physics-grounded reward models and verifiers built from simulator ground truth and robot trajectories, then distill to few-step samplers so an agent can act inside the model at interactive rates.
- Evaluation of Learned Simulators:
- Build evaluation for action-following, physical plausibility, and long-horizon drift, scored against simulator ground truth and by whether a policy trained inside the model transfers to a robot, not by FVD.
What they're asking for
- Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical fieldEducation
- Experience training generative or predictive sequence models end to end at multi-node scale (video, 3D, or latent dynamics) and fluency in PyTorch from prototype to production-scale trainingSkill
- Ability to independently design, execute, and analyze machine learning experiments, from hypothesis to ablation to conclusionSkill
- Working command of the fundamentals (diffusion and flow matching, long-context sequence modeling, and understanding of their practical limits)Skill
- Serious approach to evaluation and experience building metrics that influenced modeling decisionsSkill
- Publications or contributions to research on generative models for image, video, or 3D contentSkillPreferred
- Experience designing video tokenizers or VAEs, with understanding of trade-offs between compression and rollout fidelitySkillPreferred
- Built action-conditioned or interactive world models for games, driving, or embodied agentsSkillPreferred
- Experience with model-based reinforcement learning or planning in a learned latent spaceSkillPreferred
- Experience building streaming or causal video generation with KV caching and rolling context at interactive ratesSkillPreferred
- Worked with real robot trajectory data (LeRobot, Open X-Embodiment) and understanding of noisy action labelsSkillPreferred
- Contributions to open-source generative model projects or related infrastructureSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
MEMBER OF TECHNICAL STAFF – WORLD MODELS ABOUT US Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. RESPONSIBILITIES - Generative World Model Architecture: - Design, train, and scale action-conditioned video and latent dynamics models: diffusion transformers with rectified flow, block-causal autoregressive hybrids, causal video tokenizers whose compression ratio sets how far a rollout survives. - Action Conditioning & Latent Actions: - Condition rollouts on robot action chunks and camera trajectories, and recover pseudo-actions from unlabeled video with inverse dynamics and latent action models, so training is not capped by what teleoperation produced. - Long-Horizon Stability & Memory: - Attack compounding drift by training on the model's own rollouts (diffusion forcing, self-forcing) and carrying persistent scene state as 3D Gaussians or point maps, so minute-long rollouts stay coherent. - Post-Training & Distillation: - Fine-tune against physics-grounded reward models and verifiers built from simulator ground truth and robot trajectories, then distill to few-step samplers so an agent can act inside the model at interactive rates. - Evaluation of Learned Simulators: - Build evaluation for action-following, physical plausibility, and long-horizon drift, scored against simulator ground truth and by whether a policy trained inside the model transfers to a robot, not by FVD. REQUIREMENTS - Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field - Experience training generative or predictive sequence models end to end at multi-node scale (video, 3D, or latent dynamics) and fluency in PyTorch from prototype to production-scale training - Ability to independently design, execute, and analyze machine learning experiments, from hypothesis to ablation to conclusion - Working command of the fundamentals (diffusion and flow matching, long-context sequence modeling, and understanding of their practical limits) - Serious approach to evaluation and experience building metrics that influenced modeling decisions NICE TO HAVE - Publications or contributions to research on generative models for image, video, or 3D content - Experience designing video tokenizers or VAEs, with understanding of trade-offs between compression and rollout fidelity - Built action-conditioned or interactive world models for games, driving, or embodied agents - Experience with model-based reinforcement learning or planning in a learned latent space - Experience building streaming or causal video generation with KV caching and rolling context at interactive rates - Worked with real robot trajectory data (LeRobot, Open X-Embodiment) and understanding of noisy action labels - Contributions to open-source generative model projects or related infrastructure
Keep exploring
Available Data & Analytics roles
These current listings are available to explore now.
- Senior Manager, ConstructionWaymo · Hybrid
- Senior Machine Learning Engineer, Driving BehaviorsWaymo · Hybrid
- Senior Machine Learning Engineer, Perception LLM/VLMWaymo · On-site
- Senior Staff Technical Lead Manager, Perception OptimizationWaymo · Hybrid
- Technical Training Manager (WRA/RS)Waymo · Hybrid
- Senior Machine Learning Engineer, Runtime and ServingWaymo · Hybrid