AI Engineer
About the company
Our client is an early-stage consumer AI company in New York building a new kind of dating service — one that cuts through the noise and helps daters find someone who genuinely resonates. It's founded and led by the founder and former CEO of one of the largest dating apps in the US, with $18M raised and a top-tier venture firm on the board. A team of five today, working together in person in their Flatiron office. Unusually for a startup at this stage, they keep humane hours — roughly 9:30am to 5:30pm.
The role
You'd join as the founding AI engineer, owning the AI systems behind the core product experience — from conversational intelligence to matchmaking insight generation. You'll design the prompting, evaluation, and model infrastructure that makes the AI reliable and continuously improving, and you'll own the model serving layer in production. The role is deliberately flexible: if you're drawn to user-facing AI and prompt architecture, that's the shape it takes; if you're more excited by ML and recommendation systems, it bends that way instead. You'd report to the Head of Engineering, on an AI team of you plus part of leadership.
What you'll do
- Design and maintain the prompts, structured outputs, and orchestration systems that power the product's AI features
- Build evaluation infrastructure to measure AI quality at scale, and keep the LLM judges aligned with human judgment
- Run structured experiments across models, prompts, and configurations to optimize quality, cost, and latency
- Own the model serving layer — deployment, inference infrastructure, model versioning, and cost/latency in production
- Build internal tooling — dashboards, admin interfaces, debugging workflows — that makes model behavior visible to the team
- Apply lightweight ML and ranking models to improve match quality, and recognize when they outperform a prompt
- Fine-tune open-source models using SFT, DPO, or RLHF where it cuts cost and latency
- Design vector databases and memory architectures for personalization and context-aware experiences
- Keep the systems fair, auditable, and human-in-the-loop where it matters
What we're looking for
These are hard requirements. Check each one that applies to you.
Don't check every box? No worries — submit your info here and we'll keep you in mind for roles that could be a better fit.
Submit Your InfoNice to have
- Experience with voice AI systems, or with recommendation and ranking systems
- Consumer-facing AI/LLM product experience at real scale
- Background at an AI research lab or a foundation-model company
- Experience with recommendation and personalization engines
- Hands-on with Langfuse, Braintrust, TensorZero, or Promptfoo
Tech stack
Python, FastAPI, SQLAlchemy, PostgreSQL, and Langfuse, plus evaluation tooling such as TensorZero and Promptfoo.
Before you apply
- This is an in-person role in the NYC (Flatiron) office. Screening asks directly whether you're based in the NYC area or willing to relocate.
- The team wants someone who has genuinely built on top of non-deterministic LLM output — eval suites, LLM-judges, and the eval-to-prompt iteration loop. A strong ML background with no applied LLM work isn't the profile.
- The process is a phone screen, a 45-minute technical screen, then a 3-hour on-site: resume narrative, culture, and 1.5–2 hours of live coding plus real-world systems design. No Leetcode.
Compensation & logistics
- $200k–$300k base + competitive equity
- New York City (Flatiron) — on-site with the team
- Roughly 9:30am–5:30pm — genuine work-life balance for an early-stage startup
- Open to visa transfers (for example OPT or H-1B transfers); no net-new visa sponsorship
- Full-time · hiring 1 for this role