All roles
AI Engineer
Ship the models and pipelines behind the insights.
Stuttgart, Germany — hybrid (2–3 days/week in office)Full-time / Working student (Werkstudent)Engineering
The analysis engine — sentiment, relationship scoring, milestone detection, the AI-written insights — is the heart of the product. You'll own how we call LLMs and classical models: the prompts, the chunking, the evaluation, the cost and latency. Privacy is a hard constraint: chat content is processed transiently and never stored.
What you'll do
- Own the LLM pipeline (currently OpenAI + Gemini via a provider abstraction) — prompts, chunking, structured-output schemas, retries, and cost control.
- Build the evaluation harness so a prompt or model change is a measurable improvement, not a vibe.
- Improve the non-LLM analytics — response-time modelling, conversation segmentation, mood trends — for accuracy and speed.
- Keep inference cheap and fast enough to run at consumer scale; profile and cut token spend.
- Uphold the privacy model: transient processing, aggregate-only telemetry, no chat content leaving the request.
What we're looking for
- 3+ years building production software, with real experience shipping LLM- or ML-backed features (not just calling an API once).
- Strong TypeScript / Node.js (our stack) or the ability to get there fast from another typed language.
- You've built prompt/response pipelines with structured outputs and know how they fail.
- Comfortable with evaluation: building datasets, defining metrics, catching regressions.
- English at C1+.
Nice to have
- NLP background — embeddings, classification, sentiment.
- Experience with cost/latency optimisation for LLM apps (caching, batching, model routing).
- Interest in on-device / in-browser inference.
- German.
Languages
English at C1 or above is required. German is a strong plus.
Working student
As a working student: ~16–20 h/week, paired with a senior engineer, owning a bounded piece of the pipeline (e.g. the eval harness or one analysis module).