Hey Bob · voice-bridge · built 2026-08-15 · barge-in 2026-08-19 · the picture page — the README carries the words
Voice entry into your own LLM tools: shortcut → speech → the right session gets the message or gets born → the result comes back as voice — and the same shortcut stops it talking when you have heard enough. This page renders the model as built: architecture, four example runs, the router's working, the session lifecycle, and what each build round got wrong first. Numbers marked measured come from the acceptance runs; the rest are from the 2026-08-14 spike.
1 · decisions
Omnigent is the execution platform; the own component is one small router CLI — with no target registry and no session typology: the router only addresses, interpretation belongs to the existing skill layer.
Session pool over real Claude Code / Codex CLIs. API: spawn, message, SSE, list, approval resolve. Stop is non-sticky: a stopped session "sleeps"; one message revives it — resume is the base primitive, not a feature.
bob route — two questions, nothing morePer utterance: (1) affinity — does a pooled session hold the context this needs? → continue there. (2) placement — else, where should the new session be born (project repo or home)? No targets.yaml, no session types. Default yolo; external side effects stay gated in skills/CLIs (R-7).
What "fetch this week's episodes" means is understood by the session — via the global skill fleet's USE WHEN routing (S-11), later enriched by the Wienerdog digest. The router never reinvents it: it doesn't understand the domain, it addresses.
bobsay + the spoken logThe session speaks its result (prosody → ElevenLabs, per-call voice, playback lock). Playback runs sentence by sentence with prefetch, so first audio comes sooner and an interruption is sentence-precise. One Swift / AVAudioEngine child per call keeps the sound output open across sentences. Only its heard acknowledgement advances the append-only log (grouped daily) — at once the recency source (who spoke last, derived at read time), routing metadata, and searchable chronicle.
No always-on ear (ABS-3 intact): hold Globe+Ctrl+Alt → record → Scribe STT → bob dictate; releasing any key of the chord sends. No wake word, and pauses cost nothing — recording follows the chord, not the sound. The same chord is the interrupt — there is no second shortcut to learn.
Pressing while Bob talks kills that answer (and the queued chunks of the same answer), writes the unheard half down, and keeps everything else silent until the ack. The next utterance is presumed to belong to the session that was cut, which is told where the cut fell. Hush touches audio, never a session — killing playback is honest because the ledger is written after playback, so the unplayed part was never a fact.
The Omnigent server is a daemon — exempted from the no-daemon rule as user-launched, localhost-only (R-15). As an adopted component it gets an integration contract, like Wienerdog.
2 · architecture
bobsay and its owned Swift player. The TypeScript code fetches and trims sentence audio; the native player keeps one output session open until the call ends. A dataPlayedBack acknowledgement, not merely scheduling audio, permits each spoken-log append. Missing native builds warn and use afplay; an ElevenLabs failure falls back to macOS say. The press is also the interrupt (iteration 4): before the recorder starts, bob hush kills whatever is playing, appends what went unheard to logs/interruptions.jsonl, and drops a pause marker into state/playback/ so nothing else speaks until the acknowledgement — the router's ack is the one voice exempt from it, because it is the event the pause is waiting for. Two logs, no pointer file: the spoken ledger records who spoke, the decision log records who was dispatched to, and "most recent interaction" is derived from both at read time. The pool has no session types: only birth context (cwd) and accumulated conversation; a sleeping session is as addressable as a running one. The only standing processes are the Omnigent server and the host daemon (R-15).3 · scenarios
S/A · new work · placement decision
"Hey Bob, check whether yesterday's episode came in."
No session holds relevant context — the router starts a new one and only decides where it should be born.
bob route.~1–2 s
ROUTERAffinity: none (no recent interaction, no match). Placement: the request points at the podcast ingester → {action: new, cwd: ~/dev/podcast-ingest}. Validates that the directory exists, dispatches, then speaks the ack: "Looking into it."~4 s decidePOST /v1/sessions (yolo) + the request via POST /events. The CLI exits here.<1 s
OMNIGENTThe host daemon launches the real Claude Code CLI in tmux in the podcast-ingest directory — it boots with the project's context.~3 s (measured)
SESSIONThe session resolves the intent (project CLAUDE.md / skills) and drives your own ingest CLI — yolo, no approval stalls.task-dependent
BOBSAY"The episode is in, but the show notes aren't up yet." — plus a line into the spoken log.+2 s
If the request pointed nowhere, the session would be born in home (~/bob) and the global skill fleet's USE WHEN routing would resolve it — the router never understands the domain, it addresses. Measurement correction from the spike: fresh sessions aren't "cold" — the 117 s once mistaken for cold start was approval wait.
S/B · affinity · to the last speaker
"Then put the show notes on the processing list too."
The podcast-ingest session spoke three minutes ago. The request is uninterpretable on its own ("the show notes"? "too"?) — the context lives in the session, not in the router.
{action: continue, session: podcast-ingest}.~4 s
ROUTERPOST /v1/sessions/{id}/events — the message lands in the session's context.<1 s
SESSIONIt knows exactly which episode — sets up the show-notes watch.warm loop: 2–3 s + work
BOBSAY"Done — the show-notes watcher checks hourly."+2 s
The "and tell me when it's up" half is proactive notification (nudging) — phase two: the same bobsay/CHANNEL path, just fired from a cron job.
S/C · affinity · back to an older session
"What happened to this morning's mac-configuration run?"
Started from a terminal this morning (via omnigent claude, so the pool sees it). Other sessions have spoken since; its process may already be stopped — doesn't matter: a sleeping session is addressable.
{action: continue, session: …}. Beyond the candidate window it would take a lookup_ledger round first, and the continue would be tagged reachback.~4 sThe same path reaches arbitrarily far back — explicit reach-back has no horizon (real precedent: returning to a July session to provenance-check a suspicious derived datum). MVP limit: the pool only sees sessions started through Omnigent — hence terminal work is worth starting with omnigent claude.
S/D · barge-in · you have heard enough
"Bocs, inkább csak annyit mondj, hogy összesen hány rész volt."
A session is reading out a long answer — twelve series, episode by episode. Somewhere in the second sentence you already know you asked for the wrong thing.
interruptions.jsonl, starting at the sentence that was playing: a half-heard sentence counts as unheard.<1 s
PTTRecords while you speak. Nothing else may play — a queued answer from another session waits.your pace
ROUTERSees a fresh barge-in (inside the follow-up window, newer than the last dispatch, to a session still in the pool) and is told to strongly prefer continuing that session — unless the utterance clearly addresses something else. The dispatched message carries a bracketed note naming where the cut fell.~4 s
ROUTERSpeaks the ack — exempt from the pause, because the ack is what the pause is waiting for — then lifts the marker in a finally. The ack is expendable; the un-muting is not.+2 s
BOBSAYThe cut session answers the new question — and, judging from its own answer text, decides whether the unheard half still matters: "Negyven rész." Then the other session's queued message finally plays.+2 s
Esc while holding cancels instead: no dispatch, and bob resume lifts the pause at once, so queued speech does not wait out the marker's 180 s deadline. That deadline exists only so a crashed roundtrip cannot mute the machine — and when it fires, it says so on stderr. Design: voice-barge-in-plan.md.
4 · detail
bob route CLI worksA per-invocation command with file-based state — the D9 four layers. No watching process: it doesn't even follow SSE, because speaking the result is the session's job (bobsay).
Stage 3 may run up to three times: the model can ask once for a transcript peek and once for a ledger reach-back. A second reach-back request becomes the deterministic fallback — it is asking for context it has already been given. A second "torn" answer is different: it buys no further peek, but the decision it arrived with is validated and stands, because it is a usable answer rather than a request. Stage 4 is separate from the schema on purpose — a decision can be perfectly well-formed and still name a session that does not exist. It caught exactly that on the first live run.
| Tier | Source | Why it's cheap |
|---|---|---|
| 1 · pool metadata | GET /v1/sessions: server-generated title, workspace, status (running/idle/waiting — a session waiting on a question is a likely target of "yes, do it"), recency | discovered, platform-maintained — zero own machinery |
| 2 · the two event logs | bobsay appends session id + spoken line after every playback; the router appends every decision with its dispatch target. Recency is derived from both at read time — no pointer file. | byproduct of speaking and of routing; the spoken sentence is exactly router-prompt-sized, and a derived value cannot go stale |
| 3 · targeted peek | only when the model says it is torn: the end of both transcripts (order=desc), fetched once | on demand, over a shortlist of two — never a corpus grep per utterance |
| + reach-back | the model asks for it by name (lookup_ledger): the full spoken ledger, every day, scored by how much of the query a line covers | only when the utterance refers to work the candidate window has hidden — the one path with no time horizon |
Hardening trigger pre-registered (R-4): if tier 3 becomes frequent or slow, an FTS index over transcript extracts — born with a reindex regenerator.
A discriminated union, not one object with optional fields — so a decision that is structurally valid but impossible to carry out cannot be expressed in the first place. There is no shape in which continue exists without a session id.
{"action":"continue", "session_id":"…", "request":"…", "ack":"…"}
{"action":"new", "cwd":"~/dev/podcast-ingest", "request":"…", "ack":"…"}
{"action":"clarify", "question":"…"}
{"action":"lookup_ledger", "query":"…"} // reach-back: grep the whole ledger, ask again
// any of the first three may add:
"candidates": [{"session_id":"…", "reason":"…"}] // torn → peek at those two
One headless call produces it — llmp claude -p, stripped to the bone with --setting-sources "", --allowed-tools "" and --strict-mcp-config: 2,548 input tokens instead of 45,914, and no personal CLAUDE.md competing with a strict JSON contract. This call thinks; it does not act. The deliberation order is fixed in the prompt, and every decision lands in a JSONL log — auditable why it routed where.
Nothing here can crash into silence. A model that times out, answers nonsense, names a session that does not exist, or a dispatch that fails — all land on the same deterministic fallback: a spoken question, logged with its reason, exit 0. Never a guess.
| # | Question | Signal | Decision |
|---|---|---|---|
| 1 | affinity | follow-up phrasing + the most recent interaction inside the window | continue there (S/B) |
| 2 | affinity | content match with pool titles / spoken lines | continue there — wakes sleepers too (S/C) |
| 3 | affinity | torn between two candidates | peek at the end of both transcripts, then decide — once per utterance |
| 4 | affinity | refers to past work, no candidate matches | lookup_ledger — grep the whole ledger, no horizon, then decide — once per utterance |
| 5 | placement | the request names or implies a project | new session in ~/dev/<name> (S/A) |
| 6 | placement | no concrete place | new session in home — the skill layer resolves intent |
| 7 | — | uninterpretable | clarify — asks back in voice, never guesses |
bobsayA tiny CLI that sessions call (allowlist: Bash(bobsay:*)): clean-for-speech (markdown out) → prosody (emphasis, pauses) → ElevenLabs with the requested voice (fallback: macOS say) → sentence audio files → one Swift / AVAudioEngine player for the call (lock, so sessions don't talk over each other) → heard acknowledgement → a log line into the daily spoken/*.jsonl. The log plays three roles: recency source (who spoke last — derived at read time, no pointer file) · routing metadata · searchable chronicle. The sessions' prompt convention asks for the one-sentence speakable summary; if silent finishes recur, the responsibility moves into a Claude Code Stop hook (the PAI-proven pattern).
The player starts lazily and stays alive across all sentence files in this invocation. It exits when the call finishes or is interrupted; it is not a daemon. The next sentence is still prefetched, and a partially heard sentence still counts as unheard. Build the helper with bun run build:player. Without it, the compatibility path warns and uses afplay, with longer gaps.
| Layer | Concretely here |
|---|---|
| Data | defaults.yaml · spoken/*.jsonl · logs/route-decisions.jsonl · logs/gc.jsonl · logs/interruptions.jsonl — files in ~/bob. The ledgers are deliberately not git-versioned (2026-08-19): a log is already the chronological record, and the repo has a remote now, so only the authored files (CLAUDE.md, defaults.yaml) are tracked; the transcripts are backed up as files. Reach-backs are a flag on decisions, not a fourth log. |
| CLI | bob route · doctor · log · gc and bobsay — deterministic, --json throughout, uniform exit codes (0 done · 1 could not · 2 wrong call) |
| Client | Hammerspoon PTT (hold-to-talk → bob dictate), with the Raycast script + Monologue as typed fallback; later CHANNEL (Telegram) over the same CLI |
| Headless | the router's decision call (claude -p — through llmp when present — capable model first, isolated from personal settings) |
5 · lifecycle
The key primitive, spike-verified: Omnigent's stop is non-sticky — the process dies, the conversation survives, and the next message relaunches automatically. "Closing" is therefore not a hard state but two cheap, reversible moves.
Verified on a real sweep, not on the spike's word for it: bob gc stopped five idle sessions, a message revived one in 4.1 s with its transcript intact, and S/C then reached a gc-stopped session by content alone and got an answer from its full context.
bob gc (launchd daily or manual); since revival is a few seconds, aggressive stopping is risk-free hygiene. It stops and never deletes, and that asymmetry is the whole safety argument: a wrong stop costs four seconds, a wrong delete costs the record a reach-back might want in July. Sessions that are running or waiting are never touched however long they have been at it — a long task is not idle, and a session parked on a question is waiting for a person.bob logs reach-backs, and a recurring pattern is a hardening trigger; (b) provenance check — legitimate use: the transcript is the audit trail of agent work (the git-history analogue), no limit warranted.6 · decisions · closed 2026-08-15 (D10)
bob route. MVP-2 (when it chafes): own PTT — built 2026-08-16 as Hammerspoon hold-to-talk + ElevenLabs Scribe (the key already in the Keychain; Whisper would have needed an account the machine lacks).src/contracts/README.md carry their essential sentences.7 · as built · 2026-08-15
Four things the design had right in spirit and wrong in detail. Each was found by something running, not by re-reading the plan — and in each case the design was corrected rather than the implementation bent around it.
The prompt listed project names; the only absolute path in sight was the home directory, so the model composed <home>/hey-bob. No session was born in the wrong place — the check caught it and the fallback spoke. The prompt now states the root and the exact shape <root>/<name>.
A reach-back failed although the model did ask to search: it queried "conference badge printing July" against a line containing everything but "July", and the grep demanded every word. Ledger search now scores partial matches and ranks by coverage.
"Spawned sessions inherit the machine's auth chain" holds for what the host daemon spawns. Launched from Raycast or launchd, plain claude is simply not logged in — so the first utterance from there fell back, and deterministically so would every one after: there are no credentials to retry with. The router now uses llmp claude, which carries its own credentials.
The worse half was that it failed quietly: the deterministic fallback is built not to shout, so the bridge degraded politely and nothing surfaced until someone read the log. bob doctor now exercises the real decision call — the failure appears at the check instead of at the microphone.
| Change | Why it was unavoidable |
|---|---|
Decision log gains target_session_id | The recency derivation needs to know who a dispatch reached. For a new decision that session does not exist until the API answers, so it cannot be read off the decision — it has to be written down. |
| The ack is spoken in-process | Same speak() code, same playback lock, same ledger, one less process spawn. Still sessionless, so the recency derivation ignores it. |
~/dev stayed a constant | C3's fields are enumerated in the plan and this is not one of them. Making it configurable would be a contract change, not a convenience. |
bob doctor grew to eight checks | Config, C6 convention, server, loopback bind, host daemon, agent, the router's own call, and a full spawn round trip in a throwaway session it deletes afterwards. |
Run through the Raycast script in a bare environment — the script leg, entered as text. The hotkey-and-dictation leg in front of it is the one part still unexercised, because it belongs to the owner.
| Step | Result | Measured | Target |
|---|---|---|---|
| S/A | new session in ~/dev/website, spoke its own answer | ack 7.4 s, result +16 s | ack ≤ ~5 s |
| S/B | continued that session; it resolved "those" from its own context | 17.1 s to speech | ≤ ~10 s |
| S/C | revived a gc-stopped session by content, answered from full context | 28.6 s to speech | ≤ ~10 s |
| nonsense | spoken question, pool untouched (8 → 8 sessions) | — | — |
The targets are missed and the breakdown says why: of the ~7 s acknowledgement, roughly 4 s is the routing decision, 2 s is actually speaking the sentence (playback is awaited on purpose, so the ledger can never claim something that was not heard), and ~1 s is process start. S/B and S/C are dominated by the session's own turn. This is exactly the evidence D10 deferred the model tier to — router_model started capable deliberately, and a faster tier would take ~4 s out of every acknowledgement. A week of decision-log data decides it, not one afternoon.
8 · as built · barge-in · 2026-08-18/19
The barge-in went in as five increments, each verified before the next. Nothing below was found by re-reading the plan; every one was found by something running, and two of them only by a person's ears. They share a shape worth carrying forward: every signal this feature sends lands somewhere as a fact about the world, and the receiver acts on it.
Two tickets minted in the same millisecond sort by pid, and pid order is not arrival order — so a later ticket could sort ahead of one already playing, and both judged themselves holder. Pre-existing, invisible until playback got busy enough. The queue now decides who may try; an exclusive-create marker decides who holds.
An amendment took the C6 text past Omnigent's 4,096-character limit for a terminal_launch_args entry. Every createSession came back 400, the router fell back, and a perfectly clear spoken request was answered with "I did not understand". Now the ceiling and a budget below it are constants, a test fails while the sentence is still in the editor, and bob doctor reports the length.
The marker went down after the kill, as the plan sketched — but a killed holder releases its ticket instantly and the next waiter polls every 25 ms. Four seconds of another session's queued speech drove through that gap, over the person mid-press. The window runs from the press, so it goes down before anything is killed.
The ack may speak through the quiet window — but it always arrives after the speech the window is holding back, so plain FIFO put it behind a ticket that could not move until the window lifted, and the window lifts only once the ack is spoken. Three minutes of silence, ended by the safety deadline. The exempt ticket now overtakes the queue it exists to release; the atomic marker was always the real exclusion, so nothing overlaps.
A take too quiet to transcribe returned before route ever ran — so nothing lifted the pause the press had opened, and the queue stayed muted for the marker's full three minutes. Every non-routing path now lifts it, and the dispatch script fires bob resume after dictate returns. Whoever opens the quiet window owes it a closer on every path; a deadline is a safety net, not a closer.
The cut bobsay died by SIGTERM, so the session saw exit 143 — a failed command — and did the sensible thing: it ran it again, in full, five seconds later, before the follow-up had even arrived. What the person heard as "it started reading the list again" was a retry. A deliberate interruption now exits 0 and says so, because 0 is the strongest "do not retry" an agent understands.
| Change | Why |
|---|---|
| Sentence chunking landed early, not last | The C6 convention makes most answers a single bobsay call, so call-level precision would report "nothing was heard" on nearly every barge-in — the cut-point metadata would be worthless exactly where it matters most. |
| The clip's own dead air had to go | With stitching in place a small gap remained between sentences; measurement put it inside the audio — 0.28–1.12 s of trailing silence per clip, ended by a codec blip that defeats the one-pass trim. Two deterministic ffmpeg passes now cut it, keeping a short breath. |
| An exchange that reached no session lost its id | The prompt offered [xN] ids for every exchange while the contract accepted only dispatched ones — so a request repeated after a clarify read as a correction of that clarify, was refused, and produced another clarify. Never offer an address that cannot be used. |
bob doctor grew to ten checks | The two additions are the quiet window (a pause past its deadline is a failure, naming bob resume) and the C6 length budget. |
9 · playback update · 2026-09-26
Sentence prefetch hid synthesis time, but starting and draining a new afplay process for every sentence added silence. The inbox app's Mac worker also used launchd's Background scheduling class, which added further delay. The shared native player removes repeated device startup; the user-requested inbox playback now uses Interactive scheduling.
The trim keeps a 200 ms sentence breath. System-audio recordings of the same complete six-sentence text measured mean gaps of 0.317 s directly and 0.284 s from the app, down from 1.437 s and 1.687 s. These are measured runs, not a latency guarantee. The first sentence still needs synthesis and device startup; slow synthesis can still exhaust prefetch.
The spoken ledger, interruption record and resume boundary are unchanged: a sentence is heard only after the player's completion acknowledgement. SIGINT/SIGTERM kills the owned player, and the helper watches for parent death. See the diagnosis and verification report for tests and recording methodology.