Senior AI Engineer, Agents (Python) (f/m/d)
Job Description & Overview
Bliro:
The AI assistant for field sales
Takes desk work off sales reps' plates entirely, with more than one million customer touchpoints documented
Trusted by lighthouses of the German Mittelstand like igus, Lanxess and Brillux
Backed by 468 Capital, Lea Partners, Rockstart and more, with over €3M raised
See Bliro in action: https://bliro.io
Why We're Hiring This Role:
Our chat agent and our voice to voice phone agent process and document thousands of customer interactions a week for field sales teams. Every answer has to be right, fast and useful on the road.
Our backend engineers build the platform. We need a senior AI engineer who owns agent quality: what the agents say, how well they say it, and how we prove they got better.
You'll own prompting and context, model benchmarking for our use cases, evals, and the agent harness behind both agents. You join a small team of high performing specialists who own the core of our agents. This role is about orchestrating and evaluating models, not training them.
What You'll Do:
30 Days: Map Bliro's chat and voice agents end to end. Partner with the backend engineers, review real user conversations, identify the biggest failure modes, and ship an initial improvement.
60 Days: Own agent quality measurement. Build eval sets and metrics for our key use cases, benchmark models on quality, latency and cost, and ship a change with a measurable improvement against the baseline.
90 Days: Establish a repeatable model for improving agent quality. Run evals on every prompt, model and harness change, turn real conversations into new test cases automatically, and set the longer term roadmap for agent quality.
Who You Are:
Evidence of ownership and shipping LLM agents to real users and proving with data that they got better. You can demonstrate excellence and pursued a technical major at one of the major technical universities like TUM, ETH, KIT, RWTH, EPFL, etc.
Strong Python and the ability to debug agent behavior across prompts, tools, context and model choice.
Hands on experience building eval pipelines, LLM judges and model benchmarks, with enough statistics to tell a real gain from noise.
A habit of reading raw conversations and turning what you