Jeevan Surya Maddu

WhatsApp Family Agent

A tool-using LLM agent my family uses without ever seeing a form, and the decision to stop

Solo · 2026

The WhatsApp agent logging bottle feeds and sleep from free-text messages sent by the family

TL;DR

Tracking a newborn's feeds and diapers through an app meant opening it and tapping through forms, which nobody reliably did. So I built an agent that lives in WhatsApp. Any family member sends a normal message in whatever words they would use, the agent works out what happened, and the entry appears in the tracking app. Six skills, 125 tests, real non-technical users. Then the habit lapsed and I decided to stop rather than defend it.

The problem

The friction was not the tracking, it was the interface. A form is fine for the person who installed the app and hopeless for everyone else in the house, and a log with gaps in it is close to worthless. WhatsApp was already where the family talks, so the right move was to meet them there and accept whatever phrasing arrived: 11am 90ml breast milk, or 90ml, or something considerably less structured than either.

What I owned

What I built: the six skills and their schemas, the deterministic parsing path, the local journal, the per-person clarifying questions, the 125 tests, and the container setup it runs in. What I reused: WhatsApp as the interface, a hosted LLM for the tool calling, and the protocol work underneath, which is someone else's: py-huckleberry-api by Woyken, MIT licensed, credited in the repo.

The hard part

Two decisions carried it. Regex first, model second: most real messages match a handful of patterns, and those go down a deterministic path that costs nothing, returns instantly, and keeps working when a model provider is down. The LLM handles only the tail, through tool calling, and when a required field is missing or confidence is low the agent asks instead of guessing, because a wrong volume in the log is worse than one extra question.

And designing for a backend that can break. The tracking app has no official API, so the client underneath is reverse-engineered and can break without warning. Every parsed intent is written to a local journal before anything is dispatched, so a backend failure costs a replay rather than a lost entry. Multiple senders write into one shared household account, so the journal records who sent what and any pending question is tracked per person, or one parent's answer would resolve a question the agent asked the other.

Outcome

It worked and the family used it. Then the tracking habit faded, the paid subscription underneath stopped being worth its cost, and the honest read was that the product had lost its reason to exist. What it proved is worth keeping: WhatsApp is where the family already is, and the lowest-friction place to log anything. A household runs on a growing pile of single-purpose apps, and each one is another thing to install, learn and remember to open. So rather than defend the original, I am turning it into a family centre: one conversation where the household logs expenses, habits and whatever else it needs to keep on top of, and manages them regularly, without adding another app. That direction is also better for the thing this project still lacks, which is evaluation. A mis-logged nap is an annoyance, so correctness here is soft and easy to avoid defining. A wrongly parsed amount or a double-counted transaction is unambiguously wrong, which forces a definition of correct, which is what an eval set actually is.