“How do you stop a model from being confidently wrong?
- How do you stop a model from being confidently wrong?
- How do you digitize a workflow without killing the judgment that made it work?
- What's the MVP for a tool someone will actually open twice?
- How do you decide which requests get the expensive model?
- How do you design for a user whose worst day is your average Tuesday?
- How do we ship public-interest software on a non-profit budget?
Not a prototype.
A system in production.
One platform I built and operate end to end — the figures below are measured, not modeled.
- Comments processed
- 5.4M+
- Tokens per cycle
- 1.4B
- Production services
- 19
- Self-hosted GPUs
- 8×
- Criminal statutes
- 12
- Min. evaluation-gate agreement
- 90%
- Lower inference cost
- ~65%
- Of measured community reach
- 81.6%
26,000+ posts monitored
~€0.0035 per unit, blended
Owned end to end
A100 / H100 / H200, multi-GPU inference
Structured legal evaluation, not keywords
Benchmarked vs. a certified EU dispute-settlement body
vs. running a frontier model on everything
Identified via network/graph discovery
Systems I build.
Not a skills cloud — the categories a production AI system actually needs.
- I
AI & inference
Serving and evaluating models in production, not just calling an API.
- vLLM
- Qwen3
- OpenAI / Anthropic / Gemini APIs
- Multi-GPU serving
- A100 / H100 / H200
- Tensor parallelism
- Structured / grammar-constrained decoding
- Prefix caching
- Model benchmarking & evaluation
- Inference cost optimization
- II
Data & distributed systems
The pipelines and queues that move millions of records without losing any of them.
- Python
- FastAPI
- PostgreSQL
- Redis / BullMQ
- Distributed queues
- Worker architectures
- Large-scale ingestion
- Corpus construction
- III
Production infrastructure
I own the infrastructure this system needs to stay up — deploys, rollbacks, and the pager.
- Docker
- Kubernetes
- GCP
- AWS Lambda
- Hetzner
- Monitoring & telemetry
- Deployment
- Rollback
- IV
Application engineering
The operator console that separates what the machine decided from what a human decided.
- TypeScript
- NestJS
- Next.js
- RBAC
- Audit logging
- Operator consoles
- Case review workflows
- V
Automation
Where a report has to come from a real device, not an API call.
- Appium
- REST APIs
- Physical-device orchestration
- Evidence capture
Half a career on each side
of the same gap.
I've spent the years since 2019 building production software where regulation actually bites — ERP and indirect-tax compliance, security operations, and now production AI systems for civic accountability and legaltech. Every regulated domain shares the same problem shape, and engineers who can cross between the legal, product, and infrastructure sides of it are rare.
Today I build and operate a production AI system out of Berlin — model serving, evaluation, and the infrastructure underneath — for civic accountability work in German criminal law. I don't make legal decisions. I build the systems the people who do rely on.
Selected case studies.
Problem narratives from the field — no client names, no internal tools.
- ICase Study
The Illegal Comment Problem
An end-to-end production AI system — discovery, staged classification, structured legal evaluation, and automated reporting — for enforcing German criminal law on social media at scale.
- Published 2026-03
- DE · StGB
- 8 min
Read - IICase Study
Shipping Civic Tech on a $30 Budget
What changes when your entire production stack has to fit on a 4GB box, survive a public audit, and cost less than lunch.
- Published 2026-04
- EU · NGO
- 6 min
Read - IIICase Study
Evidence That Holds
Building artifact pipelines where every record has to be traceable, timestamped, and reviewable by someone who wasn't there.
- Published 2026-04
- DE · Evidentiary
- 7 min
Read
Latest notes.
Field notes on product, process, and the places where law and software disagree.
Get in touch.
A few questions, so I know what you're working on. No newsletter, no follow-up — I read every one myself.