docs: Sprint 16 — fix kimi-k2.6:cloud latent bug across all 6 running docs

Sprint 16 (commit 25e29c7) is a 2-line fix that switches
OLLAMA_MODEL from kimi-k2.6:cloud to gpt-oss:20b and bumps
max_tokens from 800 to 4000. The Sprint 13 LLM endpoint has
returned picked_count=0 silently since 2026-06-05 because
kimi-k2.6 is a reasoning model that burns the token budget
on internal reasoning and never produces the JSON answer.
The library fill (Sprint 6+) silently took over. Discovered
while answering the user's "is there anything else to refine?"
question.

Live verification: 5/5 test weeks return picked_count 15-21
(was 0/5 before). 11/11 vitest cases pass (4 new from S16 +
7 from S14). npm run build green. No new runtime deps. No
schema change. No UI change.

This commit updates the 6 running docs that track sprints:

- .agent/plan.md — Sprint 16 section (S16.1-S16.4 + Done
  when + Out of scope) added after the Sprint 15 sections.
  Documents the diagnosis (kimi-k2.6 reasoning model), the
  fix (gpt-oss:20b + max_tokens=4000), the 4-case Vitest
  contract test, and the live verification commands.
- .agent/context.md — Sprint 16 decisions (D1-D5), open Q1,
  and file:line references added.
- Review/sprint16-verification.md — NEW: full diagnosis +
  2-line fix + 4-test contract + live verification (5/5
  test weeks return picks, table) + 5-risk table + 4
  follow-up tickets.
- Review/ui-nielsen-audit.md — Sprint 16 status block added
  after the Sprint 15 Round 3 block.
- fix-ui-audit.md — Sprint 16 section (T9.1-T9.5) added
  after the Sprint 15 section. T9.1 documents the 2-line
  fix in detail (config.py + llm_plan.py + .env). T9.5
  surfaces 3 follow-up tickets.
- Review/handoff-ui-audit.md — Batch L line in the deploy
  list, Sprint 16 section after the Sprint 15 section, TL;DR
  Sprint 16 line, Last-updated footer updated.
- docs/HANDOFF.md — Sprint 16 section after Sprint 15, Last-
  updated footer updated. Notes the corrected model choice
  and the 4 follow-up tickets.

All 6 docs now reflect Sprint 16. The Sprint 13 LLM endpoint
now works as designed. Every future "Ask the LLM" click
will actually use the LLM to pick meals from the 77-recipe
library (was silently using the library fill instead). The
_ask_llm helper is still the single F9-full seam.
This commit is contained in:
2026-06-08 07:26:22 -07:00
parent 25e29c714d
commit c54d3ffc1f
7 changed files with 269 additions and 4 deletions
+46
View File
@@ -869,3 +869,49 @@ All five go under `devDependencies`. Runtime bundle size unchanged (503.82 kB be
- DB went 67 → 77. **Library is well past the 4-week coverage threshold (77 unique vs 84 picks needed).**
- LLM test (Sprint 13, week 2026-08-03, prompt "comfort food, no repeats from past 2 weeks"): `picked_count=0 / filled_count=21 / failed_count=0`.
- No code changes; pure content op. No pre-existing WIP files touched.
---
## Sprint 16 — Fix Sprint 13 LLM-model latent bug — 🚧 IN PROGRESS
**Why this sprint:** User asked "is there anything else to refine?" Sprint 13's `/api/llm/plan` endpoint has been silently broken since 2026-06-05 — every call returned `picked_count=0` because `kimi-k2.6:cloud` is a reasoning model that burns the `max_tokens` budget on internal `reasoning` and never produces the JSON answer. The library fill (Sprint 6+) silently took over every time, masking the bug.
### T9.1 · Backend: model switch + token bump
**Files:** `backend/app/config.py:38`, `backend/app/api/llm_plan.py:117`, `backend/.env` (or `docker-compose` env).
- **Root cause:** kimi-k2.6 is a reasoning model. On the Sprint 13 prompt (47 recipes, 21 picks), it uses 8200+ chars of `reasoning` and the 800-token `max_tokens` cap finishes with `finish_reason: length` and `content=''`.
- **Fix part 1 (config.py):** `OLLAMA_MODEL: str = "gpt-oss:20b"`. gpt-oss is OpenAI's open-source 20B non-reasoning model. Same `chat/completions` endpoint, same `messages` format.
- **Fix part 2 (llm_plan.py):** `max_tokens: 4000` (was 800). 21 picks × ~100 chars + reasoning + boilerplate ≈ 2100+ chars. 4000 gives 2x headroom.
- **Fix part 3 (.env / docker-compose):** `OLLAMA_MODEL=gpt-oss:20b`. Pydantic settings read env first, so the `.env` change is what actually fixed the running container. The `config.py` default is a backup.
### T9.2 · Frontend: Vitest contract test on LLM response shape
**File:** `frontend/src/api/llm.test.ts` (NEW, ~100 lines, 4 cases).
- **Case 8a:** `mealPlannerApi.llm.plan({prompt, week_start})` POSTs to `/llm/plan` with the payload.
- **Case 8b:** `response.plan_id` is a valid UUID.
- **Case 8c:** `picked_count`, `filled_count`, `failed_count` are non-negative integers summing to ≤ 21 (one week).
- **Case 8d:** `reasoning` is string or null (handles both the success and library-fills-everything cases).
- Uses `vi.spyOn(mealPlannerApi.llm, 'plan')` to mock the call site directly (avoids the DataCloneError that came from mocking `axios.post`).
- 11/11 tests pass (4 new from S16 + 7 from S14).
### T9.3 · Sprint 16 verification gate
- [x] `npm test` — 11/11 cases pass in ~30 ms.
- [x] `npm run build` — green (bundle 503.82 kB unchanged).
- [x] Live API: 5/5 test weeks return `picked_count` 15-21 (was 0/5 before).
- [x] Backend env verified: `docker exec mealplanner-backend-1 env | grep OLLAMA_MODEL``gpt-oss:20b`.
- [x] No new runtime dependencies (no npm install).
- [x] No migration. No schema change. No UI change.
- [ ] Commit on host + push.
### T9.4 · `Review/sprint16-verification.md` (NEW)
- Full diagnosis + 2-line fix + 4-test contract + live verification (5/5 weeks return picks) + risk table + follow-up tickets.
### T9.5 · Follow-up tickets (carry forward from Sprint 15)
- Lower `_DAILY_LIMIT=140` to 45 (S15 follow-up, still pending).
- Backend test infrastructure (venv on `docker-willester` is broken).
- CI integration of Vitest tests.