AI progress: honest milestone bar from closed JSON contract units

Replace the char-count estimate (EST_OUTPUT_CHARS) with ContractProgress:
points are earned only when real output units complete — first token (5),
each closed contract key / array item (5-90), root closed (95), stored (100).
Reasoning/thinking streams no longer fake progress: the first thinking
delta fires the first-token milestone ('the model is alive') and holds.
Surface the job's tone through ProcessingJobOut/JobInfo so the bar can
label itself 'Summarizing dry wit…'. Test rewritten to pin the new
honest-thinking semantics.
This commit is contained in:
avi 2026-09-18 15:00:04 -05:00
commit 61f2124b1c
6 changed files with 197 additions and 42 deletions

View file

@ -21,9 +21,12 @@ from shonar.services.ai._llm import (
async def _collect_reply(resp: httpx.Response, ticker) -> str:
"""Reassemble the assistant reply from a (possibly streamed) response.
Feeds the running character count to [ticker] as deltas arrive so the
UI can show progress. Servers that ignored "stream": true answer with
a plain JSON body — that path is handled too."""
Feeds the accumulated content *prefix* to [ticker] as deltas arrive;
the ticker derives progress from completed JSON contract units only
(see ContractProgress). Reasoning deltas are real work but not
contract output — they stream before the JSON and intentionally do
not move the bar. Servers that ignored "stream": true answer with a
plain JSON body — that path is handled too."""
import json as _json
ctype = resp.headers.get("content-type", "")
@ -33,11 +36,10 @@ async def _collect_reply(resp: httpx.Response, ticker) -> str:
content = _json.loads(body)["choices"][0]["message"]["content"]
except (ValueError, KeyError, IndexError, TypeError) as e:
raise ProviderTransientError("LLM sent an unreadable reply.") from e
ticker(len(content or ""))
ticker(content or "")
return content or ""
parts: list[str] = []
total = 0
async for line in resp.aiter_lines():
if not line.startswith("data:"):
continue
@ -49,18 +51,23 @@ async def _collect_reply(resp: httpx.Response, ticker) -> str:
delta = chunk["choices"][0].get("delta") or {}
piece = delta.get("content") or ""
# Thinking models (qwen3 on llama.cpp) stream a long
# reasoning_content channel BEFORE any content: counting only
# content left the progress bar frozen at "waiting for model"
# for the entire run, then jumping straight to done. Reasoning
# is real work — count it toward progress (never into the
# reply itself).
# reasoning_content channel BEFORE any content. It is real
# work but not contract output: counting it made the bar
# rocket to 99 while the JSON had not even started. The
# honest signal is the JSON itself completing.
thinking = delta.get("reasoning_content") or ""
except (ValueError, KeyError, IndexError, TypeError):
continue # keep-alives / usage chunks / odd frames: not content
if piece or thinking:
if piece:
parts.append(piece)
total += len(piece) + len(thinking)
ticker(total)
ticker("".join(parts))
elif thinking:
# First thinking delta proves the model started working:
# the ticker's 5% "first token" milestone fires once (an
# empty prefix is a no-op after that). The thinking phase
# itself earns no further points — it is not contract
# output — but "the model is alive" is real information.
ticker("")
content = "".join(parts)
if not content:
raise ProviderTransientError("LLM sent an unreadable reply.")