Eventual

Building the AI data engine for any modality and scale

Hiring — 5 openYC-W22B2B -> InfrastructureEarly

What Eventual does

Every breakthrough AI application, from foundation models to autonomous vehicles, relies on processing massive volumes of images, video, and complex data. But today’s data platforms (like Databricks and Snowflake) are built on top of tools made for spreadsheet-like analytics, not the petabytes of multimodal data that power AI. As a result, teams waste months on brittle infrastructure instead of conducting research and building their core product. Eventual was founded in 2022 to solve this. Our mission is to make querying any kind of data, images, video, audio, text, as intuitive as working with tables, and powerful enough to scale to production workloads. Our open-source engine, Daft, is purpose-built for real-world AI systems: coordinating with external APIs, managing GPU clusters, and handling failures that traditional engines can’t. Daft already powers critical workloads at companies like Amazon, Mobileye, Together AI, and CloudKitchens. We’ve assembled a world-class team from Databricks, AWS, Nvidia, Pinecone, GitHub Copilot, Tesla, and more, quadrupling our size within a year. With backing from Y Combinator, Caffeinated Capital, Array.vc, and top angels from the co-founders of Databricks and Perplexity, we’re looking to double the team now. Join us—Eventual is just getting started. Please note we are looking for someone who is willing and able to come into our San Francisco office in the Mission district 4 days / week.

5 open roles

Research Engineer, Multimodal Data
San FranciscoFull-timeAny (new grads ok)Visa: US citizen/visa only
What the role involves

About Eventual Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. As a result, robotics and video-AI teams iterate on model improvement about once a week. Most of that week isn't training — it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year. Eventual was founded in 2022 to close it. Our open-source engine, <u>Daft</u>, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop. Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate. One iteration per day becomes the norm. We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next. Join our small (but powerful!) team working together 4 days/week in our SF Mission district office. Your Role As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content. Physical AI teams have raw footage, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. We change that economics: we run vision-language models over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes. You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a research engineering role — meaning you'll read papers and run experiments, but you ship to production and your work is judged by what it does for customer training runs. Key Responsibilities Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale. Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and convolutional perception models against customer datasets and benchmarks. Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical. Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs. Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering. Work directly with researchers at our partner labs — your shortest feedback loop is their next training iteration. What we look for Strong familiarity with modern vision and multimodal models — convolution nets, VLMs, VQA, embeddings — and a sense for the SOTA that's actually deployable today vs. on a leaderboard. Experience running these models at scale on r

Software Engineer, High Performance Computing
San FranciscoFull-timeAny (new grads ok)Visa: US citizen/visa only
What the role involves

About Eventual Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. Robotics and video-AI teams now lose 20-40% of their training time to dataloading alone. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year. Eventual was founded in 2022 to close it. Our open-source engine, <u>Daft</u>, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that streams curated datasets to GPUs at line rate. Saturates B200s today. Aimed at NVL72 and Vera Rubin tomorrow. We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next. Join our small (but powerful!) team working together 4 days/week in our SF Mission district office. Your Role As a Systems Engineer on the Dataloading team, you'll build the layer that turns multi-petabyte video corpora into dict[str, Tensor] already on the GPU at line rate. We work with the top labs training Physical AI on the newest generation hardware — H100, B200, GB200, NVL72, with Vera Rubin on the horizon — on billions of dollars worth of compute, in collaboration with partners that are the largest public AI companies on Earth. Our job is to keep those GPUs fed: rank-aware sampling, NVMe caching, video and sensor co-loading, random access into clips, decode pipelining. Streaming alone can already saturate a B200; the hard part is enabling the complex sampling patterns researchers actually need without giving up a single percentage point of MFU. This is a systems engineering role for someone who feels physical pain when a system is slow. You won't need GPU experience on day one — we'll uplevel you on NVL72, CUDA, and SLURM. We will need you to bring real expertise on what happens between NVMe, network, memory, and CPU, and a deep instinct for where bytes go. Key Responsibilities Design and build the video-native dataloader: rank-aware, NVMe-cached, random-access into clips, returns tensors directly to the GPU. Profile and optimize the full data path from object store → NVMe → page cache → host RAM → device RAM. Eliminate every avoidable copy and stall. Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs. Push toward Vera Rubin bandwidth requirements. Own performance benchmarks against customer baselines (custom DataLoaders, DALI, decord, LeRobot) and against our own historical numbers — regressions get caught at PR time. Partner with researchers at our partner labs to land the loader in their training stack and measure MFU end-to-end. Work cross-team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model-output ingestion path. What we look for Obsession with systems-level performance. You can recite Jeff Dean's "numbers every programmer should know" in your sleep. You eat flamegraphs for breakfast. Strong opinions on io_uring — love it or hate it, you've earned the opinion. Live and breathe Rust, C++, or C. You reach for them when it matters and you know why. Strong familiarity with operating systems — page cache, scheduling, syscalls, NUMA, memory hierarchies. A sense for where bytes

Software Engineer, New Grad
San FranciscoFull-timeAny (new grads ok)Visa: US citizen/visa only
What the role involves

About Eventual Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. As a result, robotics and video-AI teams iterate on model improvement about once a week. Most of that week isn't training — it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year. Eventual was founded in 2022 to close it. Our open-source engine, <u>Daft</u>, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop. Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate. One iteration per day becomes the norm. We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next. Join our small (but powerful!) team working together 4 days/week in our SF Mission district office. Your Role If you're a new grad and the thought of spending the next two years bolting another LLM into another wrapper makes you tired, read on. We build real systems. We move petabytes of video through GPU clusters. We profile flamegraphs. We argue about io_uring and Parquet footers. We write Rust and Python. We talk to robotics researchers at the world's top labs about what's actually broken in their training loop, and then we go fix it. The work is hard, the feedback loops are tight, and the impact lands in customer training runs the same week you ship. As a new grad on our team, you'll work shoulder-to-shoulder with senior engineers on the layers that matter — Daft's distributed query engine, our video-native dataloader, the storage and indexing layer, the visual understanding pipeline, or the product surface researchers actually use. You'll be given real ownership early. You'll get mentorship from engineers who've spent careers on systems at AWS, Tesla, Pinecone, and Render. You'll ship code that runs on the most expensive GPU clusters on the planet. We're looking for the kind of new grad who built something obsessive in undergrad — a database, a game engine, a kernel hack, a custom decoder, a research project that wouldn't quit — not because it was assigned, but because they couldn't stop thinking about it. Key Responsibilities Contribute to features across Eventual's stack: the open-source Daft query engine, the dataloading layer, the storage and indexing layer, or the visual understanding pipeline. Profile, benchmark, and optimize real systems on real workloads — petabytes of customer video on real GPU clusters. Write clean, maintainable Python and Rust. Read papers, prototype, ship to production. Pair with senior engineers who'll teach you systems work the right way — and trust you with real ownership early. Sit in on customer calls with researchers at top Physical AI labs. The shortest path from "researcher mentions a pain" to "engineer ships a fix" is the path we want you on. What we look for Within ~1 year of graduating, or recently graduated. Strong programming fundamentals in Python, Rust, C++, or Go. Bonus poi

Software Engineer, Product
San FranciscoFull-timeAny (new grads ok)Visa: US citizen/visa only
What the role involves

About Eventual Every breakthrough Physical AI system — humanoid robots, autonomous vehicles, video generation models — is trained on petabytes of video, lidar, radar, and sensor data. But today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not the multimodal corpora that power AI. As a result, robotics and video-AI teams iterate on model improvement about once a week. Most of that week isn't training — it's finding the right data: writing CV heuristics over raw footage, paying annotators for edge cases, hand-curating clips before a cluster ever spins up. GPU bandwidth has grown 2-3× per generation. Storage and pipelines haven't. The gap widens every year. Eventual was founded in 2022 to close it. Our open-source engine, <u>Daft</u>, is the distributed data engine purpose-built for multimodal AI — already running 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at Mobileye, TogetherAI, and CloudKitchens. We are building a video-native index on top of our engine for Physical AI that collapses the data iteration loop. Describe the dataset you want, get a curated table in minutes, feed it to your GPUs at line rate. One iteration per day becomes the norm. We're building this in partnership with the top PhysicalAI labs and public AI infrastructure companies today. We have raised $30M from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and angels from the co-founders of Databricks and Perplexity. We've assembled a world-class team from AWS, Render, Pinecone and Tesla. We have spent our careers powering the last generation of PhysicalAI in self-driving, and are excited to now do this for the next. Join our small (but powerful!) team working together 4 days/week in our SF Mission district office. Your Role As a Fullstack Engineer, you'll own the product surface that researchers actually touch. Underneath us is a multimodal database indexing petabytes of video, sensors, and embeddings — but a Physical AI researcher experiences our product as the interface where they explore their corpus, write a natural-language query, watch the results come back, sanity-check clips, and ship a curated dataset to a training run. That interface is yours to design and build: visualizations over multimodal data, analytics on dataset composition, query authoring, dataset versioning, and the APIs that let customers integrate any of it into their own training stack. You'll work directly with researchers at our partner labs — your shortest feedback loop is them telling you what they wish they could see in their data. We move fast, ship to production weekly, and care more about whether researchers are actually using the surface than how clean the abstraction is. Key Responsibilities Design and build the product UI for exploring, querying, and curating multimodal datasets — including video playback, clip-level annotation, and visualizations over corpus composition. Design and build the APIs that drive the UI and that customers integrate against from their own training stacks and notebooks. Build analytics that help researchers understand their corpus: distributions over labeled axes, dataset composition over time, query result quality, training-job dataset provenance. Work closely with the Visual Understanding, Dataloading, and Storage teams so the product surface stays a thin, fast layer over a deep platform. Sit with researchers at design-partner labs, gather requirements directly, and turn them into shipped features in days — not quarters. Write high-quality, extensible, maintainable code. Take on tech debt deliberately for velocity, and pay it down deliberately when the product proves out. What we look for Fullstack engineering experience across web applications, developer-facing products, or data products. Proven track record of shipping core product features with strong user obsession, including direct collaboration with users to gathe

Software Engineer, Systems
San FranciscoFull-time3+ years$150K - $250KVisa: US citizen/visa only
What the role involves

About Eventual Every breakthrough AI application, from foundation models to autonomous vehicles, relies on processing massive volumes of images, video, and complex data. But today’s data platforms (like Databricks and Snowflake) are built on top of tools made for spreadsheet-like analytics, not the petabytes of multimodal data that power AI. As a result, teams waste months on brittle infrastructure instead of conducting research and building their core product. Eventual was founded in 2022 to solve this. Our mission is to make querying any kind of data, images, video, audio, text, as intuitive as working with tables, and powerful enough to scale to production workloads. Our open-source engine, Daft, is purpose-built for real-world AI systems: coordinating with external APIs, managing GPU clusters, and handling failures that traditional engines can’t. Daft already powers critical workloads at companies like Amazon, Mobileye, Together AI, and CloudKitchens. We’ve assembled a world-class team from Databricks, AWS, Nvidia, Pinecone, GitHub Copilot, Tesla, and more, quadrupling our size within a year. With Series A and seed funding from Felicis, CRV, Microsoft M12, Citi, Essence, Y Combinator, Caffeinated Capital, Array.vc, and top angels from the co-founders of Databricks and Perplexity, we’re looking to double the team now. Join us—Eventual is just getting started. Please note we're looking for individuals who are excited to be a part of a tight-knit team working together 4 days / week in our SF Mission district office. Your Role: As a Software Engineer on the Systems team, you will build key capabilities for the Daft distributed data engine. You will be working on core architectural design and implementation of various components in Daft. While we are an experienced team that can provide constant guidance and mentorship, we value engineers who can autonomously scope and solve difficult technical challenges. Key Responsibilities: Planning/Query Optimizer: intelligently optimize users’ workloads with modern database techniques Execution Engine: improve memory stability through the use of streaming computation and more efficient data structures Distributed Scheduler: improve Daft’s resource utilization, task scheduling and fault tolerance Storage: improve Daft integrations with modern data lake technologies such as Apache Parquet, Apache Iceberg and Delta Lake Our goal is to build the world’s best open-source distributed query engine, becoming the leading framework for data engineering and analytics. We are a young startup - so be prepared to wear many hats such as tinkering with infrastructure, talking to customers and participating heavily in the core design process of our product! What we look for: We are looking for a candidate with a strong foundation in systems programming and ideally experience with building distributed data systems or databases (e.g. Hadoop, Spark, Dask, Ray, BigQuery, PostgreSQL etc) 3+ years of experience working with distributed data systems (query planning, optimizations, workload pipelining, scheduling, networking, fault tolerance etc) Strong fundamentals in systems programming (e.g. C++, Rust, C) and Linux Familiarity and experience with cloud technologies (e.g. AWS S3 etc) Most importantly, we are looking for someone who works well in small, focused teams with fast iterations and lots of autonomy. If you are passionate, intellectually curious and excited to build the next generation of distributed data technologies, we want you on the team! Perks & Benefits In-person tight knit team with 4x a week in office Competitive comp and startup equity Catered lunches and dinners for SF employees Commuter benefit Team building events & poker nights Health, vision, and dental coverage Flexible PTO Latest Apple equipment 401k plan with match!

Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.

Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.

Questions and experiences

Nobody has asked anything about Eventual yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.

Reviewed before it appears. Do not include anything that identifies you or anyone else.

Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.

jobo is a browser extension. Open this on a computer to install it.