Research Engineer - Pre-training
Pluralis Research · Remote
MeritLog read this listing from Pluralis Research's Ashby job board and last checked it on September 10, 2026.
Source: the employer's Ashby job board. Open the original listing for current details.
Job details
- Work model
- Remote
- Salary
- Not listed by source
- Location
- USA or Australia
- Company website
- pluralis.ai
Hiring context
How this role compares at Pluralis Research
Pluralis Research has 11 live roles in MeritLog’s catalog across 3 job families, and 5 of them are in engineering. 0 of those listings publish a pay range, a disclosure rate of 0%.
Pluralis Research concentrates this hiring in:
Counted across the job boards MeritLog tracks, at the time this page was served. Pay comparisons use only listings that publish a complete range in the same currency and period.
What the role asks for
What you'd do
- Distributed pretraining: Implement and optimize model-parallel training. Data, pipeline, and tensor parallelism for large models on heterogeneous GPUs under low-bandwidth, high-latency links.
- Performance optimization: Implement techniques that reduce communication overhead while maintaining model convergence in challenging network environments.
- Elasticity and fault tolerance: Make runs survive node churn. Robust checkpointing, state synchronization, and recovery as participants join and leave.
- Run instrumentation: Build the monitoring that shows throughput, bottlenecks, and model quality across hundreds of devices.
- Hands-on distributed training (required): You've trained models across many devices in PyTorch with FSDP, DeepSpeed, Megatron, or your own implementation. You understand data, tensor, and pipeline parallelism.
- Strong engineering: Production-quality Python. Concurrency, failure handling, profiling before optimizing.
- Evidence of execution: Shipped systems, research code, open-source work, or serious personal projects.
- Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI.
What they're asking for
- Hands-on experience training or serving large language models such as Nemotron, Qwen or OLMo.SkillPreferred
- Experience with P2P networking and NAT traversal.SkillPreferred
- Experience with post-training and RL.SkillPreferred
- Experience with inference and serving systems.SkillPreferred
- Experience at proprietary, open-weight and open-source AI labsSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights (tech report https://arxiv.org/abs/2607.13332). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning https://pluralis.ai/blog/a-third-path-protocol-learning/. This setting breaks nearly every assumption of datacenter training: communication-efficient training across different parallelism axes, fault tolerance as nodes join and drop mid-run, heterogeneous compute and networks, and robustness to malicious participants. Our published methods include Subspace Networks https://arxiv.org/abs/2506.01260, Factored Gossip DiLoCo https://arxiv.org/abs/2606.22768, AsyncMesh https://arxiv.org/abs/2601.22442, and Sentinel https://arxiv.org/abs/2603.03592. As a Research Engineer you'll build the training system that takes Protocol Learning from the 8B run to frontier scale: large models on heterogeneous hardware, in physically different regions, connected by ordinary internet. KEY RESPONSIBILITIES - Distributed pretraining: Implement and optimize model-parallel training. Data, pipeline, and tensor parallelism for large models on heterogeneous GPUs under low-bandwidth, high-latency links. - Performance optimization: Implement techniques that reduce communication overhead while maintaining model convergence in challenging network environments. - Elasticity and fault tolerance: Make runs survive node churn. Robust checkpointing, state synchronization, and recovery as participants join and leave. - Run instrumentation: Build the monitoring that shows throughput, bottlenecks, and model quality across hundreds of devices. WHAT WE'RE LOOKING FOR - Hands-on distributed training (required): You've trained models across many devices in PyTorch with FSDP, DeepSpeed, Megatron, or your own implementation. You understand data, tensor, and pipeline parallelism. - Strong engineering: Production-quality Python. Concurrency, failure handling, profiling before optimizing. - Evidence of execution: Shipped systems, research code, open-source work, or serious personal projects. - Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI. NICE TO HAVE - Hands-on experience training or serving large language models such as Nemotron, Qwen or OLMo. - Experience with P2P networking and NAT traversal. - Experience with post-training and RL. - Experience with inference and serving systems. - Experience at proprietary, open-weight and open-source AI labs COMPENSATION & BENEFITS - Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary. - Remote-First Culture: Flexible work environment with team members distributed globally. - Visa Sponsorship: Optional full visa sponsorship and relocation support to either Australia or the US. - Open Problems: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones. FYI'S - We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones. - Applicants must have professional-level English proficiency (written and spoken). - Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help. We are backed by Union Square Ventures https://www.usv.com/ and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply.
Keep exploring
More Engineering roles
- Software Engineer, Fleet ManagementOpenAI · Hybrid
- Engineering Manager, Distillation & Detection PlatformOpenAI · On-site
- Forward Deployed Engineer, GovOpenAI · Hybrid
- Advanced Packaging Reliability EngineerOpenAI · Hybrid
- Engineering Manager, ArtifactsOpenAI · Hybrid
- Software Engineer, Integrity FoundationsOpenAI · On-site