Sathish Balakrishnan

I build production LLM systems — and measure them.

AI application engineer. Agentic workflows with tool calling and structured outputs on Gemini / Vertex AI, evaluation and red-teaming, and the full stack underneath (PHP 8, Python, TypeScript, GCP, Cloudflare). Everything below is live in 2026 and was built end-to-end by me.

India · ISTFull overlap with UK hours and US-East morningsRemote · EOR or contract via Media Experts LLC (US)

Get in touchRead my technical explainers

98.60%exact-match OCR on a 5,206-line gate (Tamil manuscripts)
131 → 33 → 27conversations red-teamed → faults reproduced → fixed, customer-facing assistant
1,319vouchers posted and reversed live against a real Tally company
25+tagged production releases, 450–675 smoke checks each

Work

Five systems, each with the problem, what I built, and the number that proves it.

Customer-facing LLM assistant, red-teamed

Gemini on Vertex AI · PHP 8 · MySQL · WhatsApp Business API

A sales assistant on ihayz.com that answers from the site's own content, captures leads only after explicit confirmation, and runs on a WhatsApp channel through signed webhooks.

  • Retrieval over 43 pages; confirm-before-file state machine; same-site gating and per-IP rate limits.
  • Replay harness: every recorded attack conversation re-runs after each change.
  • Cost engineering after a real overspend: thinking-token behaviour measured per call shape, per-route budgets set.

Result: 131 conversations recorded → 33 reproducible faults → 27 fixed by rules and 1 code defect, all re-verified by replay.

Cost: ~2,000 conversations a month in single-digit dollars of model fees.

Try it: the chat on ihayz.com.

Tamil manuscript OCR engine

Python · Gemini Pro / Flash on Vertex · SQLite / MariaDB · React reader

Century-old printed Tamil verse books, photographed page by page, read into a searchable, verse-bound text with a reader UI.

  • Word-level overlap crops so lines are never cut mid-word; two-model agreement with a Pro→Flash fallback on quota.
  • Quota auto-denial handling; a gate set of 5,206 lines with exact-match scoring.
  • Reader binds every line to its verse; page-scoped questions answered from the text only.

Result: 98.60% exact on the 5,206-line gate; 0 errors on a 153-line second book; 35/35 smoke checks including a real paid read.

Bank-to-Tally

Go · Python · Tally XML (port 9000) · Windows DPAPI

Bank statements (XLS/PDF) turned into classified accounting vouchers and posted into Tally ERP — the software most Indian SMEs run — with a safe reverse.

  • Go connector on the client's Windows PC; keys stored with DPAPI; REMOTEID idempotency so a re-run never double-posts.
  • Every batch reversible; balances reconciled before and after.

Result: 1,319 vouchers posted and then reversed live against a real company, balances matched to the rupee.

Ledger Core — multi-tenant admin product

PHP 8.2 API · MySQL · React / Vite / TypeScript · IMAP / SMTP

A schema-driven admin platform for small businesses: Mail with AI drafting, Tickets with intake API and dedupe, OCR Studio, Video Studio, Google Cloud, per-business module grants, and platform / brand / tenant / public-demo doors.

  • Tag-only release pipeline; a guard refuses any commit that changes behaviour without its docs.
  • Smoke suite of 452–675 checks runs on every release.

Result: 25+ tagged production releases in 2026; the mail module proven end-to-end with real customer replies.

Grounded content and video pipelines

Gemini 3.1 Pro with search grounding · Vertex image models · Gemini TTS / Chirp3 / ElevenLabs · FFmpeg on a GPU VM

Two pipelines that turn a brief into published work with gates at every step, so nothing unverified goes out.

  • Articles: draft → price/fact gate → component build → gated image prompts → renders → one-command publish (Cloudflare purge, IndexNow, sitemap, llms.txt).
  • Video: research → TTS → gated image prompts → renders → GPU render (g2-standard-32, nvenc) → publish, for multiple YouTube channels.
  • Underneath: a one-file Vertex pool — 20 projects across 4 billing accounts with rotation, cool-downs, 403/429 reason handling and a token-refresh cron.

Result: technical explainers on parameters, quantisation and the GPU memory wall published at ihayz.com/notes, every figure checked against a dated vendor page.

How I work

I have always worked for myself and for clients, never as an employee — so I am used to owning a problem end-to-end and being judged on results.

Measure before claiming

A gate set, a replay harness or a smoke suite exists before the feature is called done. The numbers on this page are the ones those tools print.

Gate the model, don't trust it

Grounded drafts still quote retired models and made-up prices. Every pipeline has a fact/price gate and a human GO before anything goes live.

Ship small, tagged, documented

Tag-based releases, docs in the same commit, backups outside the web root, restore-tested before any migration.

Own the infrastructure

GCP projects, billing and quotas, Cloudflare rules, cPanel/Apache/PHP-FPM, DNS/SSL, server migrations — I run what I build.

Cost is a feature

After one real overspend I measured thinking-token behaviour per call shape and set budgets per route. I estimate cost before any multi-call job.

Time zones

IST. Full overlap with UK working hours and US-East mornings; late-evening PST overlap by arrangement.

Stack

LLM applications

Gemini 3.x (text, image, TTS, Veo) on Vertex AI · OpenAI and Anthropic APIs · tool / function calling · structured outputs · RAG and retrieval · evaluation harnesses · red-teaming · grounding · MCP

Languages

Python · PHP 8 · TypeScript / React / Vite · Go · Bash · SQL

Data and infrastructure

MySQL / MariaDB · SQLite · REST and webhooks · GCP (Compute, IAM, billing, quotas) · Cloudflare API · Linux / cPanel / Apache / PHP-FPM · FFmpeg / nvenc · Git tag-based releases

Integrations

WhatsApp Business · Instagram Graph · LinkedIn · Tally XML · IMAP / SMTP · n8n / Make-class automation

Contact

Open to a full-time remote role (EOR: Deel / Remote.com / Oyster) or contract work invoiced through Media Experts LLC, USA. Available immediately.

[email protected]LinkedIn