Cloudglue

The video context layer for AI.

Hiring — 3 openYC-S24Early

What Cloudglue does

Cloudglue is the video context layer for AI. We make it easy for your AI to understand video. - Tinycloud - your AI agent for video, now open for beta: https://tinycloud.cloudglue.dev - Developer API Platform: https://cloudglue.dev

3 open roles

Founding Engineer, Infrastructure
San Francisco, CA, US / Remote (Seattle, WA, US; Los Angeles, CA, US)Full-time3+ years$120K - $250K0.50% - 2.50% equityVisa: US citizen/visa only
What the role involves

Cloudglue - Video Understanding Infrastructure Cloudglue is a Y Combinator-backed startup building developer APIs that turn video and audio into structured, searchable data. Think of us as the Stripe for video understanding - we handle the hard infrastructure (transcription, visual analysis, search, extraction) so developers can build on top of video without managing ML pipelines themselves. Our team has shipped large-scale systems at Snapchat and Amazon, with work presented at AWS re:Invent, KubeCon, NeurIPS, ICCV, CVPR, and DEF CON. We process millions of minutes of video for customers building search, analytics, and automation products. We’re a small, technical team where engineers have real ownership and direct impact on the product. The Role We’re looking for a founding infrastructure engineer to design and scale the backend systems that power Cloudglue’s video processing pipelines, search and retrieval infrastructure, and async job orchestration. You’ll be one of the first engineers on the team - this is a high-ownership role where you’ll shape the architecture and the engineering culture. You’ll work on: Distributed video processing and async job orchestration Search and retrieval systems across video, audio, and text ML inference serving and model pipeline orchestration Storage, indexing, and compute infrastructure for large media collections This is a systems-heavy role for someone who enjoys building reliable, high-throughput infrastructure and cares about getting the fundamentals right. What You’ll Do Distributed systems: Design and operate async, fault-tolerant job execution systems that process thousands of hours of video reliably. Search & retrieval: Build and optimize search infrastructure across video, audio, and text - including vector search, re-ranking, and hierarchical retrieval. ML infrastructure: Own the serving and orchestration layer for ML models (vision, audio, language) in production. Performance & reliability: Profile and optimize throughput, latency, and cost across large-scale video workloads. Production ownership: Build systems that are observable, well-tested, and SOC2-compatible. You’ll own what you ship. What We’re Looking For Required 3+ years of backend or infrastructure engineering experience Track record designing and operating scalable distributed systems Strong proficiency in Python, Go, and/or TypeScript Ability to reason about performance, cost, and reliability tradeoffs Nice to Have Experience with video/media processing or ML serving infrastructure Distributed systems or workflow orchestration (Temporal, Inngest, etc.) Vector search, retrieval systems, or ranking pipelines Cloud infrastructure (AWS/GCP), Docker/Kubernetes Experience with search or retrieval systems (vector databases like Milvus/Weaviate, ranking pipelines) Why Cloudglue? Video is the largest and most underutilized data source on the internet. Most software still can’t meaningfully work with it. We’re building the infrastructure to change that, and this role sits at the core of it. If you want to work on: Hard distributed systems problems with real scale Search, retrieval, and ML infrastructure that doesn’t have off-the-shelf solutions A domain (video) where the infrastructure is still being invented A small team with massive leverage …this is that role.

Amazon Web Services (AWS)GoNode.jsPostgreSQLPythonDistributed SystemsDockerServerlessSearch
Founding Engineer, Full Stack
San Francisco, CA, US / Remote (Seattle, WA, US; Los Angeles, CA, US)Full-time3+ years$120K - $250K0.50% - 2.50% equityVisa: US citizen/visa only
What the role involves

Cloudglue - Video Understanding Infrastructure Cloudglue is a Y Combinator-backed startup building developer APIs that turn video and audio into structured, searchable data. Think of us as the Stripe for video understanding - we handle the hard infrastructure (transcription, visual analysis, search, extraction) so developers can build on top of video without managing ML pipelines themselves. We process millions of minutes of video for customers building search, analytics, and automation products. The engineering problems are real: high-throughput media processing, complex search and retrieval over multimodal data, and APIs that need to be fast, reliable, and a pleasure to use. Our team has shipped large-scale systems at Snapchat and Amazon, with work presented at AWS re:Invent, KubeCon, NeurIPS, and DEF CON. We’re a small, technical team where engineers have real ownership and direct impact on the product. The Role We’re looking for a founding full stack engineer to build and ship features across Cloudglue’s APIs, developer dashboard, SDKs, and documentation. You’ll be one of the first engineers on the team - this is a high-ownership role where you’ll shape the product and the engineering culture. You’ll work across the entire stack, from React frontends to backend services to database queries. Day to day, you’ll: Own features end-to-end - from API design to UI to production deployment Work directly with founders on core product decisions Build the tools and interfaces developers use to work with large video collections Ship quickly and iterate based on real customer feedback If you like building polished developer products and moving fast across frontend, backend, and infrastructure, this role is for you. What You’ll Do Build across the stack: Design and ship features spanning frontend (React/TypeScript), backend services (Node.js/Python/Go), and databases (Postgres). Developer tools: Build and maintain REST APIs, SDKs (Python/JS), the developer dashboard, and interactive playgrounds. Product integration: Wire up ML pipeline outputs (transcription, visual analysis, search results) into usable product surfaces and API responses. Ownership: Take features from idea through implementation, deployment, monitoring, and iteration. Collaborate: Work directly with founders, infrastructure engineers, and customers. Short feedback loops, no layers of process. What We’re Looking For Required 3+ years building and shipping production web applications Strong fundamentals in both frontend and backend development Experience with API design and building developer-facing products Clear communication and a bias toward shipping Nice to Have React, TypeScript, Next.js experience Node.js backend experience SQL + query optimization Experience integrating ML model outputs into production systems Familiarity with search systems (vector databases like Milvus/Weaviate, or similar) UI/UX instincts for developer tools Why Cloudglue? Video is the largest and most underutilized data source on the internet. Most software still can’t meaningfully work with it. We’re building the infrastructure to change that - and the product surface you build is how developers will interact with it. You’ll work on genuinely hard engineering problems, ship software that real developers depend on, and have outsized influence on a product that’s defining a new category.

Amazon Web Services (AWS)Node.jsPostgreSQLReactTypeScriptServerlessPrototyping
Research Engineer
San Francisco, CA, US / Remote (Seattle, WA, US; Los Angeles, CA, US)Full-time1+ years$120K - $250K0.50% - 2.50% equityVisa: US citizen/visa only
What the role involves

Cloudglue - Video Understanding Infrastructure Cloudglue is a Y Combinator-backed startup building developer APIs that turn video and audio into structured, searchable data. We handle the hard infrastructure - transcription, visual analysis, search, extraction - so developers can build on top of video without managing ML pipelines themselves. We process millions of minutes of video for customers building search, analytics, and automation products. The research problems are real: how do you retrieve the right 10 seconds from 10,000 hours of video? How do you extract structured facts from noisy, multimodal content? How do you reason across visual and spoken information at scale? Our team has shipped large-scale systems at Snapchat and Amazon, with work presented at NeurIPS, ICCV, CVPR, KubeCon, and DEF CON. We’re a small, technical team where researchers ship code and engineers read papers. The Role We’re looking for a research engineer to work on the core multimodal retrieval and video reasoning systems that power Cloudglue. This is a 50/50 research and engineering role - you’ll design novel approaches to hard retrieval and understanding problems, and you’ll ship them into production where real customers depend on them. You’ll work across: Multimodal retrieval - finding relevant moments across visual, audio, and text signals in large video collections Structured extraction - pulling entities, facts, and relationships from video content Video reasoning - understanding temporal, causal, and semantic relationships across long-form content Evaluation and benchmarking - designing metrics and datasets to measure real-world system quality This is not a pure research role. You’ll be expected to take ideas from paper to prototype to production. But it’s also not a pure engineering role - we need someone with genuine research depth who can identify the right problems to work on and design novel solutions. What You’ll Do Multimodal retrieval: Design and improve retrieval systems that search across video, audio, and text - including embedding models, re-ranking, and hierarchical search strategies. Video understanding: Build systems that extract structured information from video - temporal segmentation, entity extraction, scene understanding, and content summarization. Model fine-tuning & integration: Fine-tune and adapt vision and language models (LoRA/PEFT, full fine-tuning) for production use cases. Evaluate open-source and proprietary models and orchestrate them in serving pipelines. Experiment and ship: Run experiments, analyze results rigorously, and turn successful research into production systems that handle real-world video at scale. Collaborate: Work directly with founders and infrastructure engineers. Short feedback loops, no layers of process. What We’re Looking For Required MS or PhD in computer science, machine learning, or a related field Research experience in one or more of: multimodal learning, information retrieval, computer vision, NLP, or video understanding Strong implementation skills in Python and PyTorch (or equivalent) Ability to independently drive research from idea to experiment to working system Nice to Have First-author publication at a top venue (NeurIPS, CVPR, ICCV, ECCV, ACL, EMNLP, SIGIR, ISMIR, ICASSP, or similar) Experience with video or multimodal foundation models (CLIP, LLaVA, Qwen3-VL, etc.) Experience with retrieval systems, embedding models, or ranking/re-ranking pipelines Experience deploying ML systems in production Familiarity with vector databases (Milvus, Weaviate) or search infrastructure Experience with model fine-tuning techniques (LoRA, PEFT, QLoRA) and training infrastructure (Ray, Kubeflow, or similar) Experience with ML inference serving (vLLM, TensorRT, Triton, or similar) Why Cloudglue? Video is the largest and most underutilized data source on the internet. Most software still can’t meaningfully search or reason over it. The research problems here - multimodal retrieval, temp

Torch/PyTorchMachine LearningReinforcement learning (RL)Natural Language ProcessingComputer VisionSearchPrompt Engineering

Roles are as last read from the company’s own listings. Openings close without notice — check the date on the listing before you spend an evening on the application.

Check the company’s own careers page — linked at the top — before a job board. A role appears there first, sometimes weeks before it is syndicated anywhere else.

Questions and experiences

Nobody has asked anything about Cloudglue yet. If you have interviewed here, what you know is worth more to the next person than anything on the rest of this page.

Reviewed before it appears. Do not include anything that identifies you or anyone else.

Company facts compiled from public sources and last refreshed 9 September 2026. Details change; treat the company’s own site as the authority.

jobo is a browser extension. Open this on a computer to install it.