Commit Graph
2 Commits
Author SHA1 Message Date
admin 25e29c714d fix(llm): Sprint 16 — switch OLLAMA_MODEL from kimi-k2.6:cloud to gpt-oss:20b
Sprint 13 (commit bae9403) set OLLAMA_MODEL=kimi-k2.6:cloud.
kimi-k2.6 is a reasoning model that burns the entire max_tokens=800
budget on internal reasoning and never produces the JSON answer
for the Sprint 13 prompt. Every /api/llm/plan call has returned
picked_count=0 since 2026-06-05. The library fill (Sprint 6+)
silently took over, masking the bug. Every "Ask the LLM" click
paid Ollama costs for nothing.

Discovered while answering the user's "is there anything else to
refine?" question. Added a temp debug log to _ask_llm, saw
raw_response='' with finish_reason: length. Verified on Ollama
Cloud: gpt-oss:20b (OpenAI's open-source 20B non-reasoning
model) returns 21 valid picks in 2074 chars on the same prompt.
finish_reason: stop. Reasoning field is 239 chars vs kimi-k2.6's
8206+ chars.

Two-line fix:
- backend/app/config.py:38 — OLLAMA_MODEL: str = "gpt-oss:20b"
  (was "kimi-k2.6:cloud")
- backend/app/api/llm_plan.py:117 — max_tokens: 4000 (was 800).
  21 picks × ~100 chars + reasoning + boilerplate ≈ 2100+ chars;
  4000 gives 2x headroom.

Plus the host's .env (or docker-compose env) was also set to
OLLAMA_MODEL=gpt-oss:20b — pydantic settings read env first, so
the .env change is what actually fixed the running container. The
config.py default is a backup for new deploys.

Plus frontend/src/api/llm.test.ts (NEW, 4 cases) — Vitest
contract test on the LLM response shape. Locks plan_id (UUID),
picked_count / filled_count / failed_count (non-negative integers
summing to ≤ 21), and reasoning (string|null). Catches
response-shape regressions so a future model swap that breaks
the JSON contract is caught at npm test time. The 4 cases: 8a
(POST to /llm/plan with payload), 8b (response.plan_id is a
valid UUID), 8c (counts are non-negative integers summing to
≤ 21), 8d (reasoning is string or null).

Verified: 11/11 vitest cases pass (4 new from S16 + 7 from S14).
npm run build green. Live API: 5/5 test weeks return picked_count
15-21 (was 0/5 before). Backend env verified:
docker exec mealplanner-backend-1 env | grep OLLAMA_MODEL →
gpt-oss:20b. No new runtime dependencies. No migration. No
schema change. No UI change.

Deploy: git pull + docker compose up -d --build backend frontend.
The .env change should already be in place; verify with
docker exec mealplanner-backend-1 env | grep OLLAMA_MODEL.
2026-06-08 07:25:37 -07:00
admin bae94037f3 feat(ui): Sprint 13 — F9-lite (Ollama Cloud free-text plan synthesis)
F9-lite reuses the pre-existing OLLAMA_* config (config.py:36-38:
OLLAMA_BASE_URL=https://ollama.com/v1, OLLAMA_API_KEY,
OLLAMA_MODEL=kimi-k2.6:cloud). Avoids the local model pull
(F9-full would be 4 GB on disk + a separate uvicorn process).
Cloud LLM — operator’s existing OLLAMA billing applies per call.

Sprint 13 splits the Sprint 11 "Generate Meal Plan" CTA into a
2-step modal: "Use the recipe library" (default, Sprint 11’s
existing flow) or "Ask the LLM" (new). The LLM path POSTs to
/api/llm/plan with a free-text prompt; the backend calls
kimi-k2.6:cloud on ollama.com, parses the LLM’s JSON picks,
creates a fresh plan, fills the LLM’s picks, and falls through
to the Sprint 6+ fillEmptySlots pattern for the slots the LLM
didn’t cover.

Backend:
- backend/app/api/llm_plan.py (NEW, ~280 lines). 1 endpoint
  (POST /api/llm/plan body {prompt, week_start}) + 4 helpers:
  - _ensure_ollama_configured — 503 on missing OLLAMA_API_KEY.
  - _serialize_library — reads up to 200 recipes for the
    family, sorted alphabetically. Cap prevents prompt-token
    overflow on kimi-k2.
  - _ask_llm — mirrors llm_matcher._ask_ollama (same URL,
    same headers, max_tokens=800, temperature=0, strips think
    blocks, 60s timeout).
  - _parse_picks — tolerant JSON parser. Handles markdown code
    fences, trailing commentary, and bare JSON. On failure
    returns []; the library fill takes over.
  - _validate_picks — drops invalid entries: missing fields,
    out-of-range day_of_week, unknown meal_type, unknown
    recipe_id. Returns a list of LLMPickedItem.
  Flow: rejects duplicate week (400) and empty library (400),
  builds the prompt, calls the LLM, validates picks, creates
  the plan, inserts the LLM-picked items, fills the rest from
  the library (Sprint 6+ pattern, re-implemented inline to
  avoid a self-HTTP-call), returns {plan_id, picked_count,
  filled_count, failed_count, reasoning}.
- backend/app/schemas/__init__.py — added LLMPlanRequest +
  LLMPlanResponse.
- backend/app/main.py:65-66 — registered llm_plan_api.router
  at the /api/llm prefix. No collision with the pre-existing
  WIP recipes.py.

Frontend:
- frontend/src/api/index.ts — added llm.plan(data) method.
- frontend/src/pages/Dashboard.tsx — added the prompt modal
  (radio for library vs. LLM + textarea for the LLM path with
  500-char counter) + new state (showPromptModal, promptMode,
  promptText, promptBusy) + extracted Sprint 11’s body into
  generateFromLibrary + added generateFromLLM. The modal is
  inline (not a separate component) because it depends on 4
  local states + 3 handlers. Click-outside-to-dismiss is
  disabled while promptBusy is true. The textarea autoFocuses
  when LLM mode is selected. Added the Button import.

LLM tolerance: a 60s timeout, parse-failure (markdown code
fences, trailing commentary), or empty response all return 0
picks; the library fill takes over. The user never sees a
crash — at worst, picked_count: 0 and the toast reads "Planned
N meals (LLM picked 0, library filled the rest)".

Verified: npm run build green (tsc 0 errors, vite 0 errors).
Bundle: 500.28 → 503.82 kB (+3.5 kB). Backend AST clean on
all 3 changed files. No new dependencies, no migration, no
pre-existing WIP files touched.

Deploy: git pull + docker compose up -d --build backend
frontend (no migration, no new dependencies).
2026-06-05 16:59:03 -07:00