Available for AI reliability diagnostics and follow-on engineering.

The model is not the product. The system around it is.

I'm an AI Systems Architect. I help founders and CTOs turn LLM prototypes into production-grade systems.

A fixed-scope diagnostic of your AI system, followed by focused engineering where it's needed. If your system works but you can't say which risks matter most, I'll tell you exactly what to fix and in what order.

PDF · no email required
01 / The problem

AI prototypes are easy. Production AI systems are harder.

  • 01Fragile prompt-wrapper architecture
  • 02Unclear LLM orchestration
  • 03Missing observability and evals
  • 04Uncontrolled latency and cost
  • 05Weak backend boundaries
  • 06Poor state, memory, and context design
02 / Services

One way in: an AI reliability diagnostic. Focused engineering after that.

AI Reliability Diagnostic

Who it's for
Technical founders and CTOs with a working AI product or prototype — under pressure from a launch, an enterprise pilot, rising cost or latency, or quality problems no one can reproduce.
What it covers
A focused review of the flows that matter: reliability, quality, cost, latency, observability, evals, retrieval, and agent behaviour. I inspect the system, not just the diagram.
Outcome
You'll know whether the system survives what's coming — the launch, the pilot, the next 10x — which three to five things to fix first, what each one costs, and what you can safely ignore. Delivered as a prioritized risk map with cost-to-fix ranges. Fixed scope and a fixed price agreed before we start, one to two weeks.

The method is public. Work through the full checklist yourself — nine areas, the discovery questions, and the audit report template I work from. If that gets you what you need, you don't need me. What it can't do is answer its own questions: the diagnostic is me reading your code, your traces, and the failure modes you actually have, rather than your architecture diagram.

The diagnostic determines which of these comes next, if any. Each is fixed scope, priced per engagement.

  • 01Cost & Latency TeardownLower spend and faster responses
  • 02Eval Harness BuildA versioned eval suite wired into delivery
  • 03Agent Reliability HardeningFewer unpredictable agent failures
  • 04Observability & Tracing SetupVisibility into what the system actually does
  • 05RAG / Retrieval Quality SprintMeasurably better retrieval and answers
  • 06AI Backend Implementation SprintHands-on build-out of agreed priorities
03 / How I think

Principles behind reliable AI systems.

  • 01Reliability is a feature. Most LLM products fail on the boring parts.
  • 02Observability before cleverness. You cannot fix what you cannot see.
  • 03Small, owned systems beat large, magical ones.
  • 04Treat prompts and evals as code: versioned, reviewed, tested.
04 / Public work

Open repositories and AI architecture notes.

Longer write-ups on how these systems are designed live on the blog. All ongoing public work is on github.com/manuelblinkert.

05 / Technical focus

LLM systems and the backends that carry them.

  • 01LLM orchestration
  • 02Backend architecture
  • 03FastAPI / Python systems
  • 04Observability and reliability
  • 05Evals and production-readiness
  • 06Knowledge infrastructure and memory systems
06 / About

A senior partner, not an agency.

Manuel Blinkert, AI Systems Architect

I'm Manuel Blinkert, co-founder and CTO of an AI product built on LLM workflows, structured memory, and async backend services — and I review other teams' systems for the failures I've had to engineer around in my own.

I work directly with founders and engineering leaders. No account managers, no hand-offs. Engagements are bounded, focused, and shipped.

07 / Contact

Let's talk about your AI system.

Have an AI system that needs a second set of senior eyes? Book a 20-minute fit call

Or start on your own: run the 10-minute self-check on your system — PDF, no email required.