
What Crustdata does
Crustdata provides live company and people data via APIs and full dataset delivery. We make hard to get data easy to use at scale. Customers use us to build their AI agent platforms and applications over our trusted data. We power builders of: -AI SDRs -AI sales platofmrs -AI recruiting platforms -AI investment/due diligence platforms -Internal sales or marketing platforms We serve use cases like: automatic pipeline building, pipeline prioritization, champion watching, company and people triggers for sales and marketing automation, investment deal sourcing We have developed technology that allows us to pipe in live data from over a dozen different data sources and deliver this data instantly to our customers. Our goal is index all the important data on the web and deliver it to customers in an easy-to-use way.
10 open roles
What the role involves
We are looking for an AI-native builder. There are limitless opportunities for you at CrustData. Depending on your skills, you may work on: Automating internal sales and customer service (CS) with AI Building demos and prototypes Solving technical questions for sales and CS Using AI to generate viral marketing contents Forward deployed in customer-facing work And more Who are we looking for: You can build things quickly You work hard You are familiar with AI and AI coding tools
What the role involves
We’re looking for a Solutions Engineering intern. This is a great opportunity to learn, take ownership, and set a foundation for rapid career growth. About Us The way information on the internet is consumed is changing. It's shifting from humans searching pre-crawled information on Google to AI agents doing real-time targeted crawling from sources of truth. But, there’s no “Google” for AI agents… yet. That’s where we come in. At Crustdata, we are building the gateway to the internet for AI agents. We already serve over 150 customers, are profitable and growing very fast. We're backed by some of the best investors in Silicon Valley including Y Combinator, General Catalyst, SV Angel, A Capital and Liquid 2 Ventures. Our mission: To be the way AI agents use and interface with the internet. What Will You Do Prototype customer solutions: build lightweight scripts and integrations to help customers go live faster Partner with AEs to design and perform demos Be the customer’s technical voice: translate customer needs into product feedback and collaborate with engineering to improve our product Drive onboarding and adoption: help enterprise clients operationalize Crustdata within their data stacks, CRMs, or AI pipelines Scale documentation & tools: build technical playbooks, templates, and guides that make every future integration faster Shape GTM strategy Who Are You Technical and want to have sales experience (SDR, AE, CS) Enjoy solving problems at the intersection of technology and sales Can communicate complex technical topics clearly and persuasively to non-technical buyers Thrive in a fast-paced, startup environment A true grinder - we work very hard People person - you love interacting with people and solving their issues Scrappy and entrepreneurial - ready to build the systems, docs, and tools from scratch. Tenacious - making sure we leave no stone unturned with the customer Systems thinker - we want to use leverage to replicate what works Note that this role is in-person in our Brooklyn (Williamsburg), New York office.
What the role involves
We’re looking to bring on our first Forward Deployed Engineer to round out our founding team. About us The way information on the internet is consumed is changing. It's shifting from humans searching pre-crawled information on Google to AI agents doing real-time targeted crawling from sources of truth. But, there’s no “Google” for AI agents… yet. That’s where we come in. At Crustdata, we are building the gateway to the internet for AI agents. We already serve over 250 customers, are profitable and growing very fast. We're backed by some of the best investors in Silicon Valley including Y Combinator, General Catalyst, SV Angel, A Capital and Liquid 2 Ventures. Our mission: To be the way AI agents use and interface with the internet. We’re looking for our first Forward Deployed Engineer. This is a great opportunity to learn, take ownership, and set a foundation for rapid career growth. What you’ll be doing You’ll be a technical partner to revenue and product teams: build rapid prototypes and demo tooling, ship customer integrations, and deliver repeatable solutions that help Sales close deals. Build rapid prototypes and POCs for prospects. Create demo tooling and reusable templates to accelerate sales. Run technical discovery and scope solutions with customers. Implement turnkey integrations (APIs, pipelines, dashboards) and hand off to CS. Connect Crustdata to customer systems and third-party tools. Iterate prototypes into repeatable reference implementations. Share feedback with Product & Engineering and write docs/playbooks. Travel occasionally for workshops or on-site work. Who you are 1+ years in customer-facing engineering (FDE, solutions, PS, pre-sales). 1+ years experience shipping production code (Python and/or JavaScript/TypeScript) + strong SQL and data stack experience (warehouses, ETL/ELT, data modeling). Strong software skills and CS fundamentals Clear communicator who can translate technical tradeoffs for business stakeholders. Comfortable working in fast, cross-functional startup environments. Why Join Impact: As the first FDE hire, you’ll have direct impact on deals and product direction. Trajectory: Learn directly from founders, YC partners, and top-tier investors Ownership: Competitive comp and meaningful equity in a profitable, fast-growing company Mission: Help build the way AI agents interface with the internet Note that this role is in-person in our SF office.
What the role involves
About the role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first 12 months. Now we need someone who can push the boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core. This is not an MLOps role. You will be researching, training, and shipping models - from paper to prototype to production. Who you are 3+ years building and shipping ML models in production — NLP, information retrieval, or entity resolution Strong with transformer architectures — you've trained and fine-tuned encoder models, not just called APIs You know how to build and evaluate retrieval systems, classifiers, and embedding models Comfortable with contrastive learning, metric learning, and representation learning Experience using LLMs for structured extraction, classification, or data generation at scale Strong Python and PyTorch A true grinder — we work very hard Founder mentality — someone who wants to be a founder in the future OR was a founder earlier What you'll be doing You'll own the ML systems that turn messy, multilingual, web-scale data into structured intelligence. Some example problems: A customer searches for "RevOps professionals" — you need to return people titled "Head of Revenue Department," "Revenue Operations Manager," and "VP Sales Operations," across English, French, and German Three different data sources list what looks like three different companies — but it's actually one. You figure out how to resolve that automatically across millions of records Given raw people data, infer the org chart — who reports to whom, what the team structure looks like, how the engineering org differs from sales Detect what technologies a company uses from unstructured signals scattered across the web Classify whether a job change was a promotion, lateral move, demotion, or just a title edit — and do it for millions of transitions Map raw job titles to canonical titles, seniority levels, and job functions — across dozens of languages and naming conventions Nice to haves Experience with entity resolution or record linkage at scale Built taxonomy or ontology systems over messy real-world data Background in multilingual NLP or cross-lingual transfer Scaled LLM inference pipelines in production Published research or open-source contributions in NLP/IR Experience with distributed training on GPU clusters
What the role involves
About the role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first 12 months. Now we need someone who can push the boundaries of what our ML systems can do. We're hiring an ML Engineer Intern to work directly with our founding team on the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core. This is a 12-week summer internship (June–August 2026). You will not be fetching coffee or watching from the sidelines. You will be researching, training, and shipping models — from paper to prototype to production. Previous interns' work has shipped to customers within weeks. Who you are Currently pursuing a Master's or PhD in Computer Science, Machine Learning, NLP, or a related field Strong fundamentals in NLP, information retrieval, or entity resolution — through coursework, research, or side projects Familiar with transformer architectures — you've trained or fine-tuned encoder models, not just called APIs Experience building retrieval systems, classifiers, or embedding models (in academic or personal projects) Exposure to contrastive learning, metric learning, or representation learning Have used LLMs for structured extraction, classification, or data generation Strong Python and PyTorch A true grinder — we work very hard Founder mentality — someone who wants to build a company someday What you'll be doing You'll own real ML problems that turn messy, multilingual, web-scale data into structured intelligence. Some example problems: A customer searches for "RevOps professionals" — you need to return people titled "Head of Revenue Department," "Revenue Operations Manager," and "VP Sales Operations," across English, French, and German Three different data sources list what looks like three different companies — but it's actually one. You figure out how to resolve that automatically across millions of records Given raw people data, infer the org chart — who reports to whom, what the team structure looks like, how the engineering org differs from sales Detect what technologies a company uses from unstructured signals scattered across the web Classify whether a job change was a promotion, lateral move, demotion, or just a title edit — and do it for millions of transitions Map raw job titles to canonical titles, seniority levels, and job functions — across dozens of languages and naming conventions Nice to haves Published research or conference papers (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.) Experience with entity resolution or record linkage at scale Built taxonomy or ontology systems over messy real-world data Background in multilingual NLP or cross-lingual transfer Open-source contributions in NLP/IR Experience with distributed training on GPU clusters Compensation & perks $8,000–$14,000/month (above market rate for SF internships) Housing stipend for those relocating to SF Direct mentorship from the founding team — no layers between you and the CEO Your work ships to production and reaches real customers
What the role involves
We’re looking to bring on our first designer to help build out our product and marketing materials. About us The way information on the internet is consumed is changing. It's shifting from humans searching pre-crawled information on Google to AI agents doing real-time targeted crawling from sources of truth. But, there’s no “Google” for AI agents… yet. That’s where we come in. At Crustdata, we are building the gateway to the internet for AI agents. We already serve over 250 customers, are profitable and growing very fast. We're backed by some of the best investors in Silicon Valley including Y Combinator, General Catalyst, and SV Angel. Our mission: To be the way AI agents use and interface with the internet. We’re looking for our first designer. This is a rare opportunity to be the first designer, shape foundational design decisions, and grow quickly alongside the company. What you’ll be doing You’ll be a core partner to the founders as we move from product-market fit to scale. You’ll own design end-to-end across product, brand, and growth: Product design Design core product experiences: dashboards, workflows, APIs surfaces, and data interactions Translate complex data and infrastructure concepts into intuitive, elegant UI Work closely with engineering to ship quickly and iteratively Design system & foundations Establish and evolve our design system (components, patterns, typography, color, motion) Set quality bars and create reusable primitives that scale with the product Brand & marketing design Define and evolve Crustdata’s visual identity Design landing pages, product pages, decks, and sales/marketing assets Create visual systems for content across web, social, and video Growth & experimentation Design assets for growth experiments (lead magnets, pages, demos, PLG flows) Support launches, announcements, and experiments with fast, high-quality design Collaborate on activation flows for our upcoming self-serve / PLG motion (B2B + B2C) Storytelling & clarity Help tell a clear, compelling story about what Crustdata does and why it matters Turn abstract concepts into visuals that “click” immediately Who you are Strong product designer - you care deeply about usability, clarity, and craft Excellent visual taste - you know what looks good Systems thinker Builder - you move fast, iterate often, and care about shipping Comfortable with ambiguity - you’re excited by zero-to-one work Collaborative - you like working directly with founders and engineers Detail-oriented but pragmatic Pluses Experience designing developer-facing or data-heavy products Familiarity with motion, interaction design, or light front-end work Designed for both B2B and B2C products Worked at an early-stage startup or as a founder Built or shipped side projects with real users Experience supporting PLG or growth experiments Why Join Impact: As the first growth hire, you’ll define how our product meets the world Trajectory: Learn directly from founders, YC partners, and top-tier investors Ownership: Competitive comp and meaningful equity in a profitable, fast-growing company Mission: Help build the way AI agents interface with the internet ***Part-time or full-time
What the role involves
About the role Skills: Python, PyTorch, NLP, LLMs, Information Retrieval, Entity Resolution, Text Classification We're building the gateway to the internet for AI agents. Our APIs already power hundreds of customers — and we went from 0 to $7M ARR in our first 12 months. Now we need someone who can push the boundaries of what our ML systems can do. We're hiring a Founding ML Engineer to own the research and engineering behind our core intelligence layer. Our platform indexes hundreds of millions of professional profiles and company records from across the web. Making that data searchable, matchable, and enriched is an ML problem at its core. This is not an MLOps role. You will be researching, training, and shipping models - from paper to prototype to production. Who you are 3+ years building and shipping ML models in production — NLP, information retrieval, or entity resolution Strong with transformer architectures — you've trained and fine-tuned encoder models, not just called APIs You know how to build and evaluate retrieval systems, classifiers, and embedding models Comfortable with contrastive learning, metric learning, and representation learning Experience using LLMs for structured extraction, classification, or data generation at scale Strong Python and PyTorch A true grinder — we work very hard Founder mentality — someone who wants to be a founder in the future OR was a founder earlier What you'll be doing You'll own the ML systems that turn messy, multilingual, web-scale data into structured intelligence. Some example problems: A customer searches for "RevOps professionals" — you need to return people titled "Head of Revenue Department," "Revenue Operations Manager," and "VP Sales Operations," across English, French, and German Three different data sources list what looks like three different companies — but it's actually one. You figure out how to resolve that automatically across millions of records Given raw people data, infer the org chart — who reports to whom, what the team structure looks like, how the engineering org differs from sales Detect what technologies a company uses from unstructured signals scattered across the web Classify whether a job change was a promotion, lateral move, demotion, or just a title edit — and do it for millions of transitions Map raw job titles to canonical titles, seniority levels, and job functions — across dozens of languages and naming conventions Nice to haves Experience with entity resolution or record linkage at scale Built taxonomy or ontology systems over messy real-world data Background in multilingual NLP or cross-lingual transfer Scaled LLM inference pipelines in production Published research or open-source contributions in NLP/IR Experience with distributed training on GPU clusters
What the role involves
The Role We are looking for a foundational member of our engineering team: a highly motivated Software Engineer to own the design, creation, and evolution of our data platform. You will be part of the team that owns the data ingestion and management infrastructure that powers Crustdata’s capabilities. If you are passionate about building robust, scalable data systems and want to see your work directly influence customers, this is the role for you. What You'll Do Architect & Build: Design, build, and maintain our core data infrastructure, including our data warehouse and data lake, using modern cloud technologies (AWS, GCP, or Azure). Pipeline Development: Develop and scale robust, fault-tolerant data pipelines (ETL/ELT) to ingest and process massive volumes of structured and unstructured data from diverse sources. Enable Data Science & ML: Create the foundational platform to support our data scientists and ML engineers. This includes building systems for feature engineering, model training, and deploying ML models into production. Orchestration at Scale: Implement and manage workflow orchestration for hundreds of daily data jobs, ensuring reliability, monitorability, and efficiency using tools like Airflow, Dagster, or Prefect. Real-time Infrastructure: Build and manage real-time data streaming pipelines using technologies like Kafka or Flink to power live dashboards and time-sensitive product features. Data Quality & Governance: Champion data quality and reliability. Implement frameworks for data validation, testing, and monitoring to ensure our data is accurate and trustworthy. Who You Are Experience: You have 3+ years of professional software engineering experience, with a significant focus on data engineering or building backend systems at scale. Strong Coder: You possess strong programming skills in Python or another modern language (e.g., Java, Go). Big Data Expertise: You have hands-on experience with modern big data technologies such as Spark, Flink, or Dask. Pipeline Orchestration: You have practical experience with workflow management tools like Temporal, Airflow, Dagster, or Prefect. Problem Solver: You are a pragmatic problem-solver who can navigate ambiguity, manage complexity, and take ownership of projects from inception to completion. Startup Mentality: You are excited to work in a fast-paced, collaborative environment and wear multiple hats. Nice to Haves Experience with real-time streaming technologies (Kafka, Pulsar, Kinesis). Familiarity with containerization and orchestration (Docker, Kubernetes). Knowledge of modern data warehousing and lakehouse architectures (e.g., Delta Lake, Iceberg).
What the role involves
We’re hiring the first engineer for our new product. You will own this from start to initial launch. Building the best APIs for realtime people and company data took us from 0 to $4M ARR in first 9 months while catering to just <1% of our customers. Now we want to democratize its access to everyone else with our new product. Who you are 5+ years of experience in shipping production software Prior experience of working with sales/gtm teams in the past would be extremely beneficial A true grinder - we work very hard Founder mentality - someone who wants to be a founder in future OR was a founder earlier and hates eating glass every day. This is a role with founder-level excitement of uncovering the unknowns minus the risks What you’ll be doing Lead the 0→1 build of a new product on top of our APIs. Including building core algorithms and architecture for creating deterministic AI agents on the fly from natural language Own product strategy and execution end-to-end: user discovery, scoping, rapid prototyping, shipping, iteration, and metrics. Collaborate with Sales/CS to define high-leverage workflows Build production-quality UX (fast, simple, correct), with pragmatic choices over perfect abstractions. Partner with the Head of Engineering on architecture, reliability, and data/privacy guardrails.
What the role involves
About Crustdata The way information on the internet is consumed is changing - from humans searching pre-crawled pages to AI agents doing real-time targeted crawling from sources of truth. But there’s no “Google” for AI agents… yet. At Crustdata, we’re building that. We serve 300+ customers and we’re growing fast. Backed by Y Combinator, General Catalyst, SV Angel, A Capital, and Liquid 2 Ventures. Our mission: be the default way AI agents interface with the internet. What you’ll do You’ll be a technical partner to revenue and product teams: build rapid prototypes and demo tooling, ship customer integrations, and turn those into repeatable solutions that help customers (and help us close deals). Build rapid prototypes and POCs for prospects Create demo tooling + reusable templates to accelerate sales Implement turnkey integrations (APIs, pipelines, dashboards) and hand off to Customer Success Connect Crustdata to customer systems + third-party tools Turn prototypes into repeatable reference implementations into the product Build AI agents for internal sales and customer success workflows Share feedback with Product & Engineering and write docs/playbooks (Optional) Join occasional on-site workshops Who you are We care way more about ability and hustle than years of experience. Undergrad (or equivalent) CS/EE/Math background Strong coding skills in Python and/or TypeScript/JavaScript Comfortable building real things end-to-end (projects, OSS, hackathons, research, prior internships - all count) Solid fundamentals + you learn fast Bonus: SQL/data curiosity (warehouses/ETL/analytics), APIs/integrations Why this internship is different You ship. You’ll work on real customer problems with real stakes. You learn fast. Small team, high trust, direct founder access. You matter. Your work will influence product direction and revenue. Future upside. Strong interns are top-of-list for full-time. If you want an internship where you ship real stuff that customers use (and you work directly with founders), this is it.
Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.
Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.
Questions and experiences
Nobody has asked anything about Crustdata yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.
Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.