Daniel Zapata
Systems running in production right now

I build AI agents that
actually ship.

I build conversational AI that takes real orders and resolves real support tickets — on WhatsApp and on the web. Not prototypes: live systems, with real customers on the other end. My doctoral research takes the same idea into cybersecurity.

See the work →
3Systems live in production
2Channels — WhatsApp & web
1Doctoral thesis in progress

Selected work

Four projects, in the order they were built. Each one solved a limit of the previous one.

01

Restaurant ordering agent on WhatsApp

Live

A customer sends a voice note saying what they want for dinner. The agent transcribes it, checks stock, builds the order, and sends back a payment link — without a human touching it.

  1. 1Someone ordersA voice note or a message on WhatsApp — however they'd say it out loud.
  2. 2It listens properlyVoice becomes text. Four messages fired off in a row get bundled into one.
  3. 3It checks realityWho the customer is, and what's actually in stock right now.
  4. 4It builds the orderAdds items, fixes changes of mind, keeps a running total.
  5. 5It takes paymentSends back a Stripe link. Nobody on our side touched it.

Try it yourself. Send it a voice note the way you'd actually order dinner — it's the live agent, not a sandbox.+52 81 8021 4366

Order on WhatsApp
The real thing. Intake on the left, the agent in the middle, its eight tools along the bottom. Click to inspect ⤢
n8nClaudeWhatsApp Business API PostgreSQLWhisperStripe
  • Eight tools the model can call — find the customer, check stock, build and update the order, generate the payment link.
  • Voice notes work. People order the way they actually talk, not by typing menu codes.
  • It waits for you to finish. Customers send four messages in a row; the agent bundles them into one turn instead of replying to each. Small detail, huge difference.

The hard part: real inventory and real payments. A demo that gets it wrong is awkward. This one charges the wrong card.

02

WhatsApp support assistant with a real knowledge base

Live

Level-1 IT support over WhatsApp for a managed service provider. It answers from the company's own documentation, walks users through fixes, and escalates to a human when it should.

  1. 1Something breaksThe user types it, records it, or photographs the error on screen.
  2. 2It reads all of itVoice notes transcribed, screenshots described. A photo of an error counts as a question.
  3. 3It looks it upSearches the company's own runbooks and price lists — not the open internet.
  4. 4It walks you throughStep by step, in plain language, and checks whether it actually worked.
  5. 5Or it gets a humanOpens the ticket, emails the team, and says so. Ransomware skips straight here.

Try it yourself. Describe an IT problem, or send a photo of an error message — it answers from the real knowledge base.+52 81 2644 6057

Message on WhatsApp
69 nodes. Every branch on the right is a decision: ticket, human handoff, email, or a satisfaction survey. Click to inspect ⤢
And how it learns. A second workflow walks Google Drive folder by folder, and turns whatever the team drops in there into something the assistant can search. Click to inspect ⤢
n8nClaudePGVector RAGOpenAI embeddingsWhisper
  • It answers from the company's own docs — runbooks, SLAs, pricing — using RAG, so it cites real procedures instead of making things up.
  • Screenshots and voice notes work too. Users send a photo of the error dialog and it reads it.
  • It knows when to stop. Suspected ransomware skips triage entirely, escalates as P1 and emails the team. Some things shouldn't be handled by a bot.

The hard part: at 69 nodes, a visual canvas can't be code-reviewed or tested. That limit is what project 03 is about.

03

The same assistant, rebuilt in Python

Live — you're on it

The support bot rewritten as a real service: FastAPI, LangChain and an embeddable web widget. Same behaviour, but now it's testable, versioned and deployable. The chat bubble on this page is this project.

  1. 1One script tagDrops the bubble onto any site. The host page's CSS can't reach in and break it.
  2. 2It streamsWords appear as they're written, instead of a spinner and then a wall of text.
  3. 3The agent decidesAnswer from the docs, ask a follow-up, or admit it doesn't know.
  4. 4Same knowledge baseThe documents from project 02 — now queried from tested code.
  5. 5Nothing gets lostTickets are written to the database first; a background worker delivers them.
💬 You can test this one right now. The bubble in the corner of this page is this exact system — same code, same knowledge base. Ask it something about my work. It speaks English and Spanish.
PythonFastAPILangChain Claude HaikuPGVectorRedis + arq DockerTraefik
  • It never loses a ticket. Requests are written to the database first and sent by a background worker, so a network blip can't drop a customer's request.
  • It refuses to guess. If the knowledge base has no good match, it says so and offers a human — instead of inventing a price.
  • One script tag to install on any site, and the host page's CSS can't break it. Rate limited per user, because every message costs model quota.

Why rewrite something that worked: version control, tests and clean rollbacks. That's the difference between an automation and a product.

04

A semi-autonomous multi-agent SOC

Doctoral research

A security analyst gets thousands of alerts a day, and almost all of them are noise. Miss the one that matters and you get a breach. My research puts a team of AI agents on that flood — with a human keeping control of anything expensive to get wrong.

AI for SecOpsMulti-agent systems Reinforcement learningPrivate on-prem LLMs MITRE ATT&CK

The idea in one sentence: stop asking the language model to do two different jobs. Let it do what it's genuinely good at — reading alerts and connecting the dots. Let a separate decision policy do what it's good at — choosing the next move when you can't see the whole picture. And keep a human on anything expensive to get wrong.

Benign — 61% Someone scanning — 22% Attacker moving inside — 12% Data being stolen — 5%

This is the whole idea. That 5% is small — but if it's real, it's a breach. A system that just picks the most likely answer says "benign" and moves on. This one holds all four possibilities at once and asks a better question: given these odds, and what each mistake would cost, what's the smartest thing to do next?

How the system is put together

Alerts arrive

Thousands a day, from the SIEM and endpoint tools. Most are noise.

It reads

A private LLM connects the dots and writes up what it thinks is happening.

It chooses

A separate policy weighs what each mistake would cost, then picks the next move.

It acts

Isolate the machine, block the account, or close it as noise.

Between choose and act sits a human. Anything expensive to get wrong waits for approval.
And it closes the loop. Whatever the analyst approves, corrects or overrides becomes the training signal for the next decision.

Why not just let the language model decide?

Because in a SOC you never actually know the true state. Is there an attacker in the network? How far in? All you ever get are noisy signals. There's a well-established way to model exactly that kind of decision — acting sensibly when you can only half-see the situation — and it's a much better fit than asking a language model to guess. That's the part I'm contributing.

The team of agents

Seven specialists rather than one do-everything model: triage sorts the flood, a threat hunter follows the attacker, intel adds outside context, containment shuts the door, forensics reconstructs what happened. An orchestrator runs the show — and a supervisor whose only job is knowing when to wake up a person.

How much rope the AI gets

Autonomy is earned, not assumed. Each level unlocks only once the system has proven itself at the one below.

LEVEL 0Watch — explains what's happening. Touches nothing.
LEVEL 1Suggest — recommends the fix; a person carries it out.
LEVEL 2Act on approval — prepares everything, waits for one click. The realistic target.
LEVEL 3Act, then tell you — with a window to undo. Only for low-risk moves it's confident about.
The part that keeps me up at night. Reward the wrong thing and you get an agent that closes real incidents — because closing tickets is what earned it points. Getting that scoring right, missed breaches versus false alarms, is the heart of the work.

Three rules I won't break

  • Nothing leaves the building. It all runs on the company's own hardware.
  • Defensive only. Blue team. No offensive tooling.
  • Every automated decision can be explained and undone. If it can't be audited, it doesn't ship.

Where the data lives

Compute in the United States, data in Mexico. Personal records never leave Mexican soil, and only two networks on the planet can reach the database at all.

US
Application · Boston VPS, Docker, behind Traefik
Traefik · TLS FastAPI · web arq · worker Redis · queue nginx · landing

Stateless by design. The containers process conversations, they don't keep them. Nothing personal sits at rest on this side of the border.

TLS Encrypted in transit.
Encrypted at rest.
MX
Database · Querétaro Azure México Central · PostgreSQL 18 + pgvector
conversations messages tickets kb_documents · embeddings

Data residency in Mexico. Every conversation, contact detail and support ticket is written to and read from a server on Mexican soil — the requirement that made this split worth the latency.

Who can reach the database

The database has no open door to the internet. The firewall answers before authentication does, so a stolen password from anywhere else is worth nothing.

✓
The application VPSOne address. The only machine that serves user traffic.
✓
My home networkFor administration, migrations and backups. Nothing else.
✕
Everything else on the internetRefused at the network layer — the connection never reaches Postgres.

Let's talk

Open to conversations about applied AI, conversational systems and research collaboration.