Member of Technical Staff, Research/RL

at Abundant — Agent simulation and RL for researchers

San Francisco, CA, US / Mountain View, CA, US / RemoteFull-time3+ years$135K - $350K0.10% - 1.00% equityYC-F24

What the role involves

About Abundant

AI models rely on two fundamental ingredients: compute and data. Abundant is building the NVIDIA of training data. NVIDIA, the leader in compute, has a peak market cap of $5T and generated $130B in revenue last year as the need for scaling compute has exploded. We believe the need to scale data is just beginning, as we move beyond SFT and human supervision to RL and Learning from Experience.

Our founding team consists of former founders, ML engineers, roboticists and data leads from Waymo, Google, Mercor and AWS. Our team has previously worked with DeepMind to deploy deep learning models at 1B user scale, trained SOTA models for self-driving at Waymo, and scaled data pipelines of tens of thousands of human annotators at YouTube. Our pioneering work in human computation, synthetic data, simulation and RL give us the advantage in delivering results to our customers.

Why now? Training data is more important and more scarce than ever before. Scaling laws dictate that linear improvement in model performance demands an exponential increase in training data. But there is only one World Wide Web and most of it has already been trained on. The next advances will require major advances in simulation, synthetic data and learning from experience.

What happens if we succeed? Abundant will be the core enabler for not only AGI, but ASI and physical intelligence. Most of the challenges in model algorithms and compute are already solved. What’s missing? The data necessary to move from general knowledge to domain expertise; from chatbots to agents; and from text to multimodal and physical AI. Ask any AI researcher or roboticist: the core bottleneck to progress is the availability of data, hence “_abundant data_”.

Abundant works with a majority of the top AI labs, as well as frontier startups and F500 enterprises.

About the Role

As a Reinforcement Learning Researcher, you will work directly with the founding team and CTO to develop our customer-facing products and internal infrastructure and tooling, as well as scale and grow a team of engineers.

As part of this role, you will lead efforts in RL, Simulation, and Data Quality, which includes running experiments on SOTA verifiers and data quality techniques, creating pipelines for data creation and curation, running data validation pipelines and experiments, and training models to validate data quality.

Requirements

Deep experience in one or more of the following

  • ML engineering or research
  • Reinforcement learning and/or agentic harnesses
  • Agent evaluation and benchmarking
  • High-scale and high-reliability batch data pipelines

We’re looking for folks that are obsessive about their work. In data, quantity is important but data quality is the differentiator for the winner in the space. In addition to this mindset, here are some skills that are a pre-requisite for working in startups:

  • Extremely clear communication. We value being able to explain very complex technical concepts in very simple terms.
  • Pragmaticsm. Able to go deep, but also to simplify and prioritize.
  • Velocity. Fast at getting work done and picking up new skills.
  • Curious about AI and technology; keeping up with the latest papers in agents, RL and benchmarks.
  • Impact-oriented. Hacker. Always works on the most important 20% of the project in successive chunks.
  • Handle extremely ambiguity. Able to unblock yourself and others.
  • Here are three different personas that will succeed at Abundant:
  • The Craftsman
  • The Craftsman cares about their work, simply for the art of it. They may put extra care into UX or design, or into data quality, or customer success. The Craftsman is energized by putting out great work.
  • The Underdog
  • The Underdog is dying to prove themselves. They’ve overlooked; they haven’t challenged by their school or company, or they are tired of politics at a FAANG company. They’re looking for a chance to maximize their full potential and talent.
  • The Antifragilist
  • Our team collectively has an uncommon trait: high pain toler

What they ask for

TensorFlowSparkTorch/PyTorchMachine LearningReinforcement learning (RL)Docker

About Abundant

Hello! 👋 We are a team of former ML engineers, founders, roboticists and ops leads who obsess about data and its impact on safe, reliable AI. We specialize in creating environments and datasets for RL by leveraging our experience in simulation and model training. By the numbers: • Powering 3 of the top 6 global AI labs and multiple Fortune 500 enterprises • Billions of training tokens generated each month, 2x month over month • Exclusive, global network of over 500 domain experts We believe humans are inherently creative, and thrive by pushing the frontier. We are working towards an abundant future--one where everyone has access to infinite intelligence, services and goods. Based in San Francisco, CA. We enjoy good food and good company. -- more info below -- Abundant is building the NVIDIA of training data. AI models rely on two fundamental ingredients: compute and data. NVIDIA, the leader in compute, has a peak market cap of $5T and generated $130B in revenue last year as the need for scaling compute has exploded. We believe the need to scale data is just beginning, as we move beyond SFT and human supervision to RL and Learning from Experience. Our founding team consists of second-time founders, ML engineers and data leads from Waymo, Google, Meta and AWS. Our team has previously collaborated with DeepMind to classify hate speech in YouTube videos, trained SOTA models for self-driving, and scaled data pipelines with thousands of human annotators. Our pioneering work in human computation, synthetic data, imitation learning and RL give us a solid advantage in delivering results to our customers. Why now? Training data is more important and more scarce than ever before. Scaling laws dictate that linear improvement in model performance demands an exponential increase in training data. But there is only one World Wide Web and most of it has already been trained on. The next advances will require new, diverse, and high-quality datasets, making training data more important and scarce than ever before. What happens if we succeed? Abundant will be the core enabler for AGI and beyond. Most of the challenges in model training are already solved. What’s missing is the data necessary to move from general knowledge to domain expertise; from chatbots to agents; and from digital intelligence to physical AI. Ask any AI researcher or roboticist: the core bottleneck to progress is the availability of data, i.e. “abundant data”. Abundant works with the most advanced AI labs and startups, as well as F500 enterprises.

Full Abundant profile

Other roles at Abundant

Similar Science, Research elsewhere

jobo is a browser extension. Open this on a computer to install it.