Files
commitea/apps/desktop/e2e/live-reginald.spec.ts
Croissant Le Doux 3de887417c feat: Reginald is real — model router + agent loop + query_project (P4)
The fixture chat panel is now a working agent. Ask Reginald a question and it
consults the real project through a tool loop, then answers in grounded prose.
Read-only v0 — writes still go through the propose-approve controls.

core (@commitea/core/agent):
- chat-client: OpenAI-wire chat completions over an injected fetch (same seam as
  gitea). Points at any OpenAI-compatible endpoint (LM Studio/Ollama/OpenAI).
- model-router: small model for prose + the read tool; big model reserved for
  later decomposition (pickModel).
- agent-loop: runAgentTurn drives call→tool→result→call until prose (or a step
  budget), recording each tool step. Injected complete + execute → fully testable.
- query-project: the single read tool's engine — compact focus/board/calibration/
  issue/search views built from scheduler + lifecycle + calibration; unbuilt views
  return a notImplemented marker (never fabricated). The model reports, never computes.
- agent-tools: query_project declaration + Reginald's system prompt.

app:
- main model bridge (model:status, model:chat) runs the loop; query_project
  reconciles the repo and builds the view. Model traffic stays in main (token/CSP).
  gitea.ts refactored to share getGiteaClient + reconcileSnapshot.
- preload + global.d.ts expose the model bridge; useChat drives the panel — real
  agent turn when a model is configured, scripted fixture reply otherwise (so
  fixture e2e is unchanged). A subtle "consulted the project" activity line.

Model config (env, defaults to LM Studio on :1234): COMMITEA_MODEL_URL /
_SMALL (google/gemma-4-e4b) / _BIG (qwen/qwen3.6-35b-a3b). COMMITEA_E2E=1 keeps
it unconfigured so the panel stays scripted.

Verified: 88 core tests green (14 agent: client parse, loop tool/error/budget,
all views) + a gated live integration test. Desktop typecheck clean, 14 fixture
e2e green. Gated live e2e drives the real app against gitea + gemma-4-e4b: asked
"what now?", Reginald called query_project and answered "focus is on issue #2"
(the real scheduler pick).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 21:23:14 -04:00

34 lines
1.6 KiB
TypeScript

import { dirname, join } from 'node:path'
import { fileURLToPath } from 'node:url'
import { _electron as electron, expect, test } from '@playwright/test'
const here = dirname(fileURLToPath(import.meta.url))
const MAIN = join(here, '..', 'out', 'main', 'index.js')
// Opt-in (GITEA_LIVE=1 + COMMITEA_MODEL_LIVE=1 + a local model on :1234). Launches
// WITHOUT COMMITEA_E2E so Reginald runs the real agent loop against the real repo.
test.describe('live Reginald', () => {
test('answers a question by consulting the real project', async () => {
test.skip(!process.env.GITEA_LIVE || !process.env.COMMITEA_MODEL_LIVE, 'live model test — opt-in')
test.setTimeout(120_000)
const app = await electron.launch({ args: [MAIN], env: { ...process.env } })
const win = await app.firstWindow()
await win.waitForLoadState('domcontentloaded')
// model configured → the live greeting + header, not the scripted demo
await expect(win.getByText('gemma-4 · local')).toBeVisible({ timeout: 20000 })
await expect(win.getByText(/I check the real board before I answer/)).toBeVisible()
const composer = win.getByPlaceholder(/Tell me what to do/)
await composer.fill('What should I work on right now?')
await composer.press('Enter')
// the agent loop ran end-to-end: it consulted the project, then answered
await expect(win.getByText(/consulted the project/)).toBeVisible({ timeout: 90_000 })
await win.screenshot({ path: join(here, '.artifacts', 'screens', 'live-reginald.png'), fullPage: true, animations: 'disabled' })
await app.close()
})
})