Research Engineer - Post-Training
Pluralis Research · Remote
This listing is no longer verified as available.
MeritLog keeps this source-backed description for reference. Availability is not verified, and there is no application link here.
Source: the employer's Ashby job board. Open the original listing for current details. Availability is not verified for this retained page.
Job details
- Work model
- Remote
- Salary
- Not listed by source
- Location
- USA or Australia
What the role asks for
What you'd do
- Build the post-training stack: You build the RL training loop end-to-end: rollout ingestion from the geo-distributed inference pipeline, reward computation, policy updates, and getting updated weights back out to the network. You set the direction, and you make things happen.
- Invent the algorithms: Standard RL recipes assume on-policy rollouts from fast, trusted hardware. You adapt them to asynchronous, high-latency, partially trusted generation: staleness tolerance, off-policy corrections, and communication-efficient policy updates.
- Ship first post-trained models: You build the evals that show the models are improving, and you take the first decentralized post-trained release from run to public artifact.
- Hands-on RL post-training: You've run RL post-training on large language models - RLHF, RLVR, or reasoning-focused RL - and touched the systems layer yourself: rollout generation, async training loops, weight synchronization. Not just launched jobs on someone else's stack.
- Strong engineering: Production-quality Python and PyTorch: concurrency, failure handling, profiling before optimizing.
- Research ability: Publications in RL post-training, asynchronous or distributed RL, or nearby fields are a strong signal. So is unpublished work you can defend in detail.
- Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI.
What they're asking for
- Experience training over slow networks, or with decentralized or federated setups.SkillPreferred
- Familiarity with serving-engine internals such as vLLM or SGLang - our rollout pipeline is a serving system.SkillPreferred
- Experience with reward modeling or building verifiable-reward datasets.SkillPreferred
- Experience with P2P networking and NAT traversal.SkillPreferred
- Experience at proprietary, open-weight and open-source AI labsSkillPreferred
Parsed by MeritLog from the employer’s own posting. The full description follows below.
Job description
Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights (tech report https://arxiv.org/abs/2607.13332). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read A Third Path: Protocol Learning https://pluralis.ai/blog/a-third-path-protocol-learning/. Agora gave us a pretrained 8B model. Post-training is how we make it useful for agentic use-cases. But every post-training stack you've seen assumes a datacenter - synchronous rollouts, fast interconnects, trusted workers. Ours gets none of that. It has to run on consumer GPUs, and Macs spread across the public internet, training a model whose weights no single participant ever holds, with rollouts arriving from a geo-distributed inference pipeline at high latencies. Your primary role is to make RL post-training work here anyway - the algorithms and the system, end-to-end. KEY RESPONSIBILITIES - Build the post-training stack: You build the RL training loop end-to-end: rollout ingestion from the geo-distributed inference pipeline, reward computation, policy updates, and getting updated weights back out to the network. You set the direction, and you make things happen. - Invent the algorithms: Standard RL recipes assume on-policy rollouts from fast, trusted hardware. You adapt them to asynchronous, high-latency, partially trusted generation: staleness tolerance, off-policy corrections, and communication-efficient policy updates. - Ship first post-trained models: You build the evals that show the models are improving, and you take the first decentralized post-trained release from run to public artifact. WHAT WE'RE LOOKING FOR - Hands-on RL post-training: You've run RL post-training on large language models - RLHF, RLVR, or reasoning-focused RL - and touched the systems layer yourself: rollout generation, async training loops, weight synchronization. Not just launched jobs on someone else's stack. - Strong engineering: Production-quality Python and PyTorch: concurrency, failure handling, profiling before optimizing. - Research ability: Publications in RL post-training, asynchronous or distributed RL, or nearby fields are a strong signal. So is unpublished work you can defend in detail. - Mission alignment: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI. NICE TO HAVE - Experience training over slow networks, or with decentralized or federated setups. - Familiarity with serving-engine internals such as vLLM or SGLang - our rollout pipeline is a serving system. - Experience with reward modeling or building verifiable-reward datasets. - Experience with P2P networking and NAT traversal. - Experience at proprietary, open-weight and open-source AI labs COMPENSATION & BENEFITS - Equity-Heavy Package: We offer significant ownership for key technical contributors in addition to a high base salary. - Remote-First Culture: Flexible work environment with team members distributed globally. - Visa Sponsorship: Optional full visa sponsorship and relocation support to either Australia or the US. - Open Problems: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones. FYI'S - We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones. - Applicants must have professional-level English proficiency (written and spoken). - Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help. We are backed by Union Square Ventures https://www.usv.com/ and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply.
Keep exploring
Available Engineering roles
These current listings are available to explore now.
- Senior AI EngineerMetropolis · Hybrid
- Senior Staff Software Engineer, AeroParkerMetropolis · Not provided by source
- Senior Software Engineer, Product FoundationsMetropolis · Hybrid
- Senior AI EngineerMetropolis · Hybrid
- Senior Central Cloud Infrastructure EngineerMetropolis · Hybrid
- Senior Software Engineer, Product FoundationsMetropolis · On-site