A non-AIError from transcribe/summarize (av.InvalidDataError on corrupt
audio, seen live with 'My recording 63') escaped the consumer, left the
row 'running', and got requeued at every engine restart forever. Both
runners now catch it, fail the row with the exception type in the
message, and keep a usable transcript (summarize completes with a
'Summary failed' note). Regression tests pin both paths (proven red
without the fix).
The reprocess endpoint poked transport_enqueue while its request session
was still uncommitted; the inline worker woke instantly, read the job row
with its own session, and summarized with the PREVIOUS click's tone
(TRACE: click Funny -> worker read 'dry, witty'; click Neutral -> ran
'funny'). The UI disabled the voice chips for the whole wrong-voice run,
which read as 'spinner does nothing until I leave and come back'.
Rows are the real queue: commit before the poke. Regression test pins
worker tone to the requested tone and Neutral clearing a prior voice;
verified red without the fix, green with it.
Backend: new ContractProgress parses the streamed summary JSON and
advances only when contract units finish — first token 5%, each closed
key 5→90 (arrays step per closed item), root close 95, stored=100.
Replaces the char-count ticker that counted reasoning chars against a
guessed 1200-char output and sat frozen at 99 for minutes. No timers,
no elapsed-time estimates; the unfinishable thinking span honestly
earns only the alive-tick. Both adapters (openai_compat, ollama) feed
the accumulated content prefix. /jobs now carries the job's tone.
App: summarize progress shows a real percentage + determinate bar with
the tone named ('Summarizing dry wit… 47% · 1:12') in both detail
widgets; indeterminate only while queued or pre-first-token.
Tests: milestone sequence verified identical for char-by-char and
chunked streaming; adapter thinking-phase test updated to the new
contract (83 passed).
Replace the char-count estimate (EST_OUTPUT_CHARS) with ContractProgress:
points are earned only when real output units complete — first token (5),
each closed contract key / array item (5-90), root closed (95), stored (100).
Reasoning/thinking streams no longer fake progress: the first thinking
delta fires the first-token milestone ('the model is alive') and holds.
Surface the job's tone through ProcessingJobOut/JobInfo so the bar can
label itself 'Summarizing dry wit…'. Test rewritten to pin the new
honest-thinking semantics.
- Engine: when the primary summarizer is out of retries (or misconfigured),
run_summarize now finishes the job on the configured rescue provider
(SHONAR_LLM_FALLBACK_*), tags the summary with the provider that wrote it,
and stores a human note in job.error; success clears stale notes.
- App (auto/lan): passes Ollama as the rescue provider when it is up.
- Detail screen: shows the rescue-swap note in plain words, a red
'Summary failed' line with Settings -> Summarizer fix instructions and a
Retry summary button on hard failure.
- Settings copy explains the fallback. 3 new pytest cases (14/14 pass);
live E2E on 2026-09-18: sarcastic summary v6 via LAN on attempt 2.
POST /reprocess?job=summarize&tone=<voice> persists the voice on the job
row (queue carries only ids, so a sweep re-enqueue keeps it), the LLM
prompt appends 'write every field in a <tone> tone — the tone colors
the wording, never the facts', and the resulting summary records its
tone (SummaryOut.tone). Plain re-summarize clears a previous tone.
Migration tone0000000001 (summaries.tone, processing_jobs.tone).
78 passed, 1 skipped; ruff clean.
Nothing read it; services/search.py already picks SQLite vs Postgres by
dialect. The 'postgres_fts' default was a leftover from the server-era
config and only invited confusion.
Registry now lists exactly two models instead of five:
- 'Whisper Base' — names the actual bundled model (was 'Base (default)').
- 'Whisper Large v3' — kept; honest copy on size/speed/RAM.
tiny/small/medium removed from the registry (validate_model_name now
rejects them; existing rows keep their saved model — all current data
is 'base', still valid).
Detail screen shows the model that WILL run by real display name
('Whisper Base (default)') instead of 'Use default (base)'; the picker
lists the two models with the default marked inline.
Backend 78 passed/1 skipped; app 34/34; engine hot-reloaded clean.
- conftest: default SHONAR_TEST_DATABASE_URL is a temp SQLite file, matching
the bundled-lite engine; set the env var to a PG URL to exercise that path
- models.py: add sqlite_where to the two partial unique indexes — without it
SQLite built a FULL unique index on recording_id (WHERE not carried over),
wrongly blocking a second export asset per recording
- test_m9: pass UUID objects (not str) to direct ORM inserts/gets; SQLite's
GUID bind processor rejects strings (asyncpg tolerated them)
Verified: pytest -q = 78 passed, 1 skipped; ruff check clean
- shared/ = portable Android-origin sources vendored from deferred/desktop-server
(app/build.gradle.kts srcDir repointed; PlaybackController.kt excluded as Android-only)
- backend/ = bundled-lite engine (SQLite + inline queue); .venv symlinked from the
old checkout, PYTHONPATH pins THIS backend's code over any editable install
- repoRoot() resolves this project dir (env SHONAR_REPO still wins); desktop-dev.sh
watches shared/ + backend/
- Verified: :app:compileKotlin + :app:test green (23 tests); engine boots on :8010,
self-migrates, /healthz ok