All open positions
Product EngineeringPS-AI

AI Engineer

Take LLM, RAG, and agent features to production quality, and turn open-ended generation quality into measurable metrics.

About the team

Develop the console and public API through which customers control their infrastructure, plus the AI features running on top.

Product Engineering develops the console and public API customers use directly, along with the AI features that run on top of them. Because we operate our own GPU servers and AI gateway (OpenGate), the infrastructure required for model serving is secured within the team, and design decisions move on a short cycle through direct discussion with the infrastructure team.

About the role

Working once in a demo and working consistently across real usage are different problems, and this role closes that gap. Response quality that degrades depending on input, generation output with no defined ground truth, and agents that fail in hard-to-predict ways are the actual problems you will handle. You decide which models, retrieval strategies, and architectures to adopt with clear reasoning, and prove the effect through metrics. We evaluate on the scope of ownership you can carry rather than years of experience.

What you'll do

  • Design and build customer-facing AI features by combining LLM, RAG, and multimodal techniques
  • Design agents that compose tools situationally and generate structured outputs
  • Design the eval framework — quantify response accuracy, consistency, and hallucination
  • Secure agent reliability with guardrails, fallbacks, and validation; build regression-prevention pipelines
  • Optimize latency and token cost, and find the right quality-cost balance
  • Build observability for response quality, latency, cost, and availability; define SLOs

What we're looking for

  • Hands-on experience designing and building LLM, RAG, or agent-based features
  • Shipped a product real users rely on by composing existing tools and models
  • Made reasoned choices on models, retrieval strategies, and prompt structure — and validated the outcome
  • Able to design services in Python or TypeScript, with solid grasp of APIs, async, and streaming
  • Chase quality regressions and incidents down to the data, then fix them structurally

Nice to have

  • Designed an LLM eval framework, or ran experiments with LLMOps tooling such as LangSmith or Langfuse
  • Solved non-obvious quality problems — RAG accuracy, fine-tuning, agent reliability
  • Handled large-scale traffic and optimized latency and token cost
  • Built and operated containers (Docker/Kubernetes) and CI/CD pipelines
  • Served models on GPU servers or self-hosted inference infrastructure

How to apply

Including the scope of work you owned and the outcomes it produced helps us evaluate your application. There are no constraints on document format.

Apply by email