AI progress: honest milestone bar from closed JSON contract units
Replace the char-count estimate (EST_OUTPUT_CHARS) with ContractProgress:
points are earned only when real output units complete — first token (5),
each closed contract key / array item (5-90), root closed (95), stored (100).
Reasoning/thinking streams no longer fake progress: the first thinking
delta fires the first-token milestone ('the model is alive') and holds.
Surface the job's tone through ProcessingJobOut/JobInfo so the bar can
label itself 'Summarizing dry wit…'. Test rewritten to pin the new
honest-thinking semantics.
This commit is contained in:
parent
6ec811393a
commit
61f2124b1c
6 changed files with 197 additions and 42 deletions
|
|
@ -21,9 +21,12 @@ from shonar.services.ai._llm import (
|
|||
async def _collect_reply(resp: httpx.Response, ticker) -> str:
|
||||
"""Reassemble the assistant reply from a (possibly streamed) response.
|
||||
|
||||
Feeds the running character count to [ticker] as deltas arrive so the
|
||||
UI can show progress. Servers that ignored "stream": true answer with
|
||||
a plain JSON body — that path is handled too."""
|
||||
Feeds the accumulated content *prefix* to [ticker] as deltas arrive;
|
||||
the ticker derives progress from completed JSON contract units only
|
||||
(see ContractProgress). Reasoning deltas are real work but not
|
||||
contract output — they stream before the JSON and intentionally do
|
||||
not move the bar. Servers that ignored "stream": true answer with a
|
||||
plain JSON body — that path is handled too."""
|
||||
import json as _json
|
||||
|
||||
ctype = resp.headers.get("content-type", "")
|
||||
|
|
@ -33,11 +36,10 @@ async def _collect_reply(resp: httpx.Response, ticker) -> str:
|
|||
content = _json.loads(body)["choices"][0]["message"]["content"]
|
||||
except (ValueError, KeyError, IndexError, TypeError) as e:
|
||||
raise ProviderTransientError("LLM sent an unreadable reply.") from e
|
||||
ticker(len(content or ""))
|
||||
ticker(content or "")
|
||||
return content or ""
|
||||
|
||||
parts: list[str] = []
|
||||
total = 0
|
||||
async for line in resp.aiter_lines():
|
||||
if not line.startswith("data:"):
|
||||
continue
|
||||
|
|
@ -49,18 +51,23 @@ async def _collect_reply(resp: httpx.Response, ticker) -> str:
|
|||
delta = chunk["choices"][0].get("delta") or {}
|
||||
piece = delta.get("content") or ""
|
||||
# Thinking models (qwen3 on llama.cpp) stream a long
|
||||
# reasoning_content channel BEFORE any content: counting only
|
||||
# content left the progress bar frozen at "waiting for model"
|
||||
# for the entire run, then jumping straight to done. Reasoning
|
||||
# is real work — count it toward progress (never into the
|
||||
# reply itself).
|
||||
# reasoning_content channel BEFORE any content. It is real
|
||||
# work but not contract output: counting it made the bar
|
||||
# rocket to 99 while the JSON had not even started. The
|
||||
# honest signal is the JSON itself completing.
|
||||
thinking = delta.get("reasoning_content") or ""
|
||||
except (ValueError, KeyError, IndexError, TypeError):
|
||||
continue # keep-alives / usage chunks / odd frames: not content
|
||||
if piece or thinking:
|
||||
if piece:
|
||||
parts.append(piece)
|
||||
total += len(piece) + len(thinking)
|
||||
ticker(total)
|
||||
ticker("".join(parts))
|
||||
elif thinking:
|
||||
# First thinking delta proves the model started working:
|
||||
# the ticker's 5% "first token" milestone fires once (an
|
||||
# empty prefix is a no-op after that). The thinking phase
|
||||
# itself earns no further points — it is not contract
|
||||
# output — but "the model is alive" is real information.
|
||||
ticker("")
|
||||
content = "".join(parts)
|
||||
if not content:
|
||||
raise ProviderTransientError("LLM sent an unreadable reply.")
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue