9 Commits

Author SHA1 Message Date
842661c9a9 Merge branch 'main' into feat/perf-pass 2026-07-09 21:54:07 +00:00
717dc7348f Merge pull request 'Memory layers: budgeted hot context (#27)' (#57) from feat/memory-layers into main
Reviewed-on: #57
2026-07-09 21:53:57 +00:00
d1a4c4410c Merge branch 'main' into feat/memory-layers 2026-07-09 21:53:51 +00:00
345b561591 Merge pull request 'query_project: implement the standup view (#28)' (#56) from feat/standup-query-view into main
Reviewed-on: #56
2026-07-09 21:53:46 +00:00
2a6322e99a Merge branch 'main' into feat/standup-query-view 2026-07-09 21:53:33 +00:00
481cacd99a Merge pull request 'SQLite cache bootstrap + single-issue mirror upsert (#3)' (#55) from feat/sqlite-cache into main
Reviewed-on: #55
2026-07-09 21:53:29 +00:00
Croissant Le Doux
f08c4935dc Performance pass: benchmark the deterministic compute path (#32)
Lock in the PLAN.md compute targets so a regression that slips an O(n²) into the
scheduler or forecast fails the suite:
- scheduler + capacity layout + Monte Carlo forecast < 1s @ 200 open issues —
  measured 232ms, comfortable headroom.
- scaling stays ~linear (400 issues ≈ 3.9x the 100-issue time; asserts < 8x to
  rule out O(n²) while tolerating jitter).

Representative fixture: 200 open issues with varied estimates/priorities/assignees
across 3 capacity lanes + a light acyclic dependency web. Bounds are the real
targets with margin so timing jitter can't flake CI; actuals are logged.

Reconcile-<5s@500 is network-bound (~2N gitea calls) and stays covered by the live
reconcile — this benchmarks the pure compute the app runs each turn. +2 core tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 15:41:56 -04:00
Croissant Le Doux
b65ee4c8ad Memory layers: budgeted hot context (#27)
Reginald's context is tiered so the model always sees what matters without ever
copying ticket data into the prompt:
- HOT (this module) — charter + active directives + the focus snapshot, packed
  under a hard token budget (2k). Assembled fresh each turn as the prompt seed.
- WARM — the append-only directive/event ledger + digest, summarized on demand.
- COLD — gitea + sidecar via query_project. Ticket bodies/comments/detail live
  here and are NEVER inlined; the model fetches them by number when needed.

- `assembleHotContext(inputs, budget=2000)`: focus (tiny, always kept) → most
  recent active directives (each while they fit ~⅔) → charter fills the true
  remainder, truncated on a line boundary. Measures the fixed tail exactly and
  reserves for header/joiner/ellipsis so the total never exceeds budget.
- `estimateTokens` (tokenizer-free ~4 chars/token, slight over-estimate so a real
  tokenizer stays under), `activeDirectives` (accepted/amended, most-recent-first).

Acceptance met: hot assembles under the 2k budget even with a ~34k-token charter;
nothing ticket-shaped is inlined (only numbers + titles for focus). +5 core tests;
full core suite + typecheck green.

Follow-up: wire assembleHotContext into the live system prompt in main (needs
charter + directives + focus at chat time) — the assembly + budget is the tested core.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 15:37:57 -04:00
Croissant Le Doux
19e83ff8ba query_project: implement the standup view (#28)
The StandupScreen already renders drift + plan + nag from real data (standup-view,
#52), but the agent's query_project standup view was a `notImplemented` stub, so
Reginald couldn't answer standup questions from deterministic data.

Implement `standupView(snap, asOf)`: today's plan (the scheduler's earliest pick
per person, with why — critical path / blocks / order), overnight drift (real
anomalies: issues sitting in review, or steeping past their estimate), and the
single stalest blocker to nag about (+ what it blocks). All deterministic; the
model narrates. Added 'standup' to the query_project tool's view enum.

Acceptance met: standup surfaces schedule drift + at least one stale blocker
(and stays calm — no nag, empty drift — when nothing is steeping). +2 core tests;
typecheck clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 15:33:20 -04:00
7 changed files with 388 additions and 5 deletions

View File

@@ -17,10 +17,11 @@ export const QUERY_PROJECT_TOOL: ToolDecl = {
properties: { properties: {
view: { view: {
type: 'string', type: 'string',
enum: ['focus', 'board', 'calibration', 'issue', 'search'], enum: ['focus', 'board', 'calibration', 'issue', 'search', 'standup'],
description: description:
'focus = Now/Next/Later; board = issues by lifecycle column; calibration = estimate-vs-actual; ' + 'focus = Now/Next/Later; board = issues by lifecycle column; calibration = estimate-vs-actual; ' +
'issue = one issue (needs filters.issueId); search = issues matching filters.query.', 'issue = one issue (needs filters.issueId); search = issues matching filters.query; ' +
"standup = today's plan + overnight drift + the stalest blocker.",
}, },
filters: { filters: {
type: 'object', type: 'object',

View File

@@ -0,0 +1,71 @@
import { describe, expect, it } from 'vitest'
import type { DirectiveEntry } from '../directives/record-directive-v0.js'
import type { Focus, ScheduledItem } from '../scheduler/scheduler-v0.js'
import {
activeDirectives,
assembleHotContext,
estimateTokens,
HOT_CONTEXT_BUDGET_TOKENS,
} from './memory-v0.js'
const item = (number: number, title: string): ScheduledItem =>
({ number, title, labels: [], order: 0, startDay: 0, endDay: 1, durationDays: 1, blockedBy: [], blocks: [], critical: false, rationale: '' })
const focus: Focus = { now: item(7, 'Fix lifecycle inference'), next: item(8, 'Webhook listener'), later: null }
const directive = (over: Partial<DirectiveEntry>): DirectiveEntry => ({
id: 'd1',
ts: '2026-02-10T09:00:00Z',
status: 'accepted',
kind: 'note',
quote: 'pilots come first',
...over,
})
describe('memory-v0 (#27)', () => {
it('estimateTokens is a slight over-estimate (~4 chars/token)', () => {
expect(estimateTokens('')).toBe(0)
expect(estimateTokens('abcd')).toBe(1)
expect(estimateTokens('a'.repeat(4001))).toBe(1001)
})
it('activeDirectives keeps accepted/amended, most-recent-first', () => {
const ds = [
directive({ id: 'a', ts: '2026-02-01T00:00:00Z', status: 'accepted', quote: 'old' }),
directive({ id: 'b', ts: '2026-02-11T00:00:00Z', status: 'amended', quote: 'new' }),
directive({ id: 'c', ts: '2026-02-12T00:00:00Z', status: 'withdrawn', quote: 'gone' }),
directive({ id: 'd', ts: '2026-02-09T00:00:00Z', status: 'proposed', quote: 'maybe' }),
]
expect(activeDirectives(ds).map((d) => d.quote)).toEqual(['new', 'old'])
})
it('assembles hot context under the 2k budget even with a huge charter', () => {
const huge = 'Charter line that goes on and on. '.repeat(2000) // ~34k tokens
const ds = Array.from({ length: 50 }, (_, i) =>
directive({ id: `d${i}`, ts: `2026-02-${String((i % 27) + 1).padStart(2, '0')}T00:00:00Z`, quote: `directive number ${i}` }),
)
const out = assembleHotContext({ charter: huge, directives: ds, focus })
expect(estimateTokens(out)).toBeLessThanOrEqual(HOT_CONTEXT_BUDGET_TOKENS)
// focus (tiny) is always kept; the charter is the part that gets truncated
expect(out).toContain('## Focus')
expect(out).toContain('#7 Fix lifecycle inference')
expect(out).toContain('…') // charter was clamped
// at least some recent directives survived
expect(out).toContain('## Active directives')
})
it('never inlines ticket bodies — only numbers + titles appear for focus', () => {
// The assembler takes no issue bodies by construction; focus shows #number title only.
const out = assembleHotContext({ charter: 'Ship the beta.', directives: [directive({})], focus })
expect(out).toContain('Now: #7 Fix lifecycle inference')
expect(out).toContain('[note] pilots come first')
expect(out).not.toMatch(/body|description|comment/i)
})
it('degrades to just focus when there is no charter or directives', () => {
const out = assembleHotContext({ charter: '', directives: [], focus })
expect(out).toBe(['## Focus', 'Now: #7 Fix lifecycle inference', 'Next: #8 Webhook listener', 'Later: —'].join('\n'))
})
})

View File

@@ -0,0 +1,102 @@
/**
* Memory layers, v0 (#27). Reginald's context is tiered so the model always sees
* what matters without ever copying ticket data into the prompt:
*
* - HOT (this module) — charter + active directives + the focus snapshot, packed
* under a hard token budget. Assembled fresh each turn; it's the system-prompt seed.
* - WARM — the append-only directive/event ledger + periodic digest, in pm-state.
* Not inlined; summarized on demand.
* - COLD — gitea + the sidecar, reached through `query_project` tools. Ticket bodies,
* comments, and per-issue detail live here and are NEVER copied into memory —
* the model fetches them by number when it needs them.
*
* The invariant: HOT stays under budget, and nothing ticket-shaped is inlined.
*/
import type { DirectiveEntry } from '../directives/record-directive-v0.js'
import type { Focus } from '../scheduler/scheduler-v0.js'
/** The hot layer's hard ceiling (#27: hot context assembles under 2k tokens). */
export const HOT_CONTEXT_BUDGET_TOKENS = 2000
/**
* Tokenizer-free estimate (~4 chars/token). Deliberately a slight over-estimate so
* a real tokenizer never exceeds what this predicts — the budget stays safe.
*/
export function estimateTokens(text: string): number {
return Math.ceil(text.length / 4)
}
/** Directives that still bind: accepted or amended, most-recent-first. */
export function activeDirectives(all: DirectiveEntry[]): DirectiveEntry[] {
return all
.filter((d) => d.status === 'accepted' || d.status === 'amended')
.slice()
.sort((a, b) => (a.ts < b.ts ? 1 : a.ts > b.ts ? -1 : 0))
}
export interface HotContextInputs {
/** The project charter markdown (hot-memory seed). */
charter: string
/** The directive ledger (any status — filtered to active here). */
directives: DirectiveEntry[]
/** The current Now/Next/Later focus, or null when nothing is scheduled. */
focus: Focus | null
}
function focusBlock(focus: Focus | null): string {
if (!focus) return ''
const slot = (label: string, item: Focus['now']) => (item ? `${label}: #${item.number} ${item.title}` : `${label}: —`)
return ['## Focus', slot('Now', focus.now), slot('Next', focus.next), slot('Later', focus.later)].join('\n')
}
function directivesBlock(directives: DirectiveEntry[]): string[] {
// one compact line each; the verbatim quote is the payload, kind is the tag
return directives.map((d) => `- [${d.kind}] ${d.quote}`)
}
/** Truncate to a token budget on a whitespace boundary, with an ellipsis marker. */
function clampToTokens(text: string, budgetTokens: number): string {
if (estimateTokens(text) <= budgetTokens) return text
const maxChars = Math.max(0, budgetTokens * 4 - 1)
const cut = text.slice(0, maxChars)
const lastBreak = cut.lastIndexOf('\n')
return `${(lastBreak > maxChars * 0.6 ? cut.slice(0, lastBreak) : cut).trimEnd()}\n…`
}
/**
* Assemble the HOT context under `budget` tokens. Priority when space is tight:
* the focus snapshot (tiny, always kept) → the most recent active directives
* (each while they fit) → the charter fills whatever budget remains (truncated).
* Never inlines ticket bodies — only charter text, directive quotes, and focus
* titles, all authored/short. Returns a single prompt-ready block.
*/
export function assembleHotContext(inputs: HotContextInputs, budget = HOT_CONTEXT_BUDGET_TOKENS): string {
const focus = focusBlock(inputs.focus)
const focusTokens = focus ? estimateTokens(focus) : 0
// fit the most recent active directives into ~⅔ of what's left after focus
const active = activeDirectives(inputs.directives)
const directiveCap = Math.max(0, Math.floor((budget - focusTokens) * (2 / 3)))
const keptDirectives: string[] = []
let directiveTokens = 0
for (const line of directivesBlock(active)) {
const t = estimateTokens(line) + 1
if (directiveTokens + t > directiveCap) break
keptDirectives.push(line)
directiveTokens += t
}
const directives = keptDirectives.length ? ['## Active directives', ...keptDirectives].join('\n') : ''
// Measure the fixed tail (directives + focus, with their joiner) exactly, then
// give the charter the true remainder — reserving for the "## Charter" header,
// the block joiner, and the truncation ellipsis so the total never exceeds budget.
const tail = [directives, focus].filter(Boolean).join('\n\n')
const tailTokens = tail ? estimateTokens(tail) : 0
const reserve = estimateTokens(`## Charter\n${tail ? '\n\n' : ''}\n…`)
const charterBudget = Math.max(0, budget - tailTokens - reserve)
const charterBody = inputs.charter.trim() ? clampToTokens(inputs.charter.trim(), charterBudget) : ''
const charter = charterBody ? `## Charter\n${charterBody}` : ''
return [charter, tail].filter(Boolean).join('\n\n')
}

View File

@@ -0,0 +1,74 @@
import { describe, expect, it } from 'vitest'
import { extractLabelFacts } from '../labels/label-schema.js'
import type { LifecycleEvent } from '../lifecycle/lifecycle-v0.js'
import type { GiteaIssue } from '../gitea/types.js'
import { buildProjectView, type ProjectSnapshot } from './query-project.js'
const asOf = new Date('2026-02-12T00:00:00Z')
function issue(over: Partial<GiteaIssue>): GiteaIssue {
const labels = over.labels ?? ['est/2d']
return {
number: 1,
title: '#1',
body: '',
state: 'open',
labels,
facts: extractLabelFacts(labels),
milestone: null,
assignee: null,
assignees: [],
createdAt: '2026-02-02T09:00:00Z',
updatedAt: '2026-02-02T09:00:00Z',
closedAt: null,
url: '',
...over,
}
}
describe('buildProjectView: standup (#28)', () => {
it('surfaces schedule drift + a stale blocker, with a per-person plan', () => {
// #7 started work 6 working days ago (steeping) and blocks #8; est is 2d → past estimate.
const steeping = issue({ number: 7, labels: ['est/2d', 'p/1'], assignee: 'christian' })
const blocked = issue({ number: 8, labels: ['est/3d', 'p/2'], assignee: 'stephen' })
const timelines: Record<number, LifecycleEvent[]> = {
7: [{ type: 'commit', at: '2026-02-04T09:00:00Z' }],
}
const snap: ProjectSnapshot = { issues: [steeping, blocked], timelines, deps: [{ issue: 8, dependsOn: 7 }] }
const v = buildProjectView('standup', undefined, snap, asOf) as {
plan: { who: string; issue: number; why: string }[]
drift: { issue: number; note: string }[]
nag: { issue: number; steepingDays: number; blocks: number[] } | null
}
// a stale blocker is nagged, and it's the steeping one that blocks another
expect(v.nag).not.toBeNull()
expect(v.nag!.issue).toBe(7)
expect(v.nag!.steepingDays).toBeGreaterThan(0)
expect(v.nag!.blocks).toContain(8)
// drift caught the past-estimate steep
expect(v.drift.some((d) => d.issue === 7)).toBe(true)
// plan gives an earliest pick per person (both assignees represented)
expect(v.plan.map((p) => p.who)).toEqual(expect.arrayContaining(['christian', 'stephen']))
// #7 leads its person's plan on the critical path
expect(v.plan.find((p) => p.issue === 7)?.why).toContain('critical')
})
it('is calm when nothing is steeping (no nag, empty drift)', () => {
const snap: ProjectSnapshot = {
issues: [issue({ number: 1, labels: ['est/2d'] })],
timelines: {},
deps: [],
}
const v = buildProjectView('standup', undefined, snap, asOf) as {
drift: unknown[]
nag: unknown | null
}
expect(v.nag).toBeNull()
expect(v.drift).toEqual([])
})
})

View File

@@ -4,9 +4,9 @@
* deterministic code (scheduler, lifecycle inference, calibration); the model * deterministic code (scheduler, lifecycle inference, calibration); the model
* only requests a shape and narrates it — it never computes (decisions.md). * only requests a shape and narrates it — it never computes (decisions.md).
* *
* v0 serves focus / board / calibration / issue. The remaining views * Serves focus / board / calibration / issue / search / standup. The remaining
* (milestone / runway / standup / search) return a `notImplemented` marker so * views (milestone / runway) return a `notImplemented` marker so the model
* the model degrades honestly instead of inventing data. * degrades honestly instead of inventing data.
*/ */
import { fitCalibration, calibrationSamples } from '../calibration/calibration-v0.js' import { fitCalibration, calibrationSamples } from '../calibration/calibration-v0.js'
@@ -121,6 +121,62 @@ function searchView(snap: ProjectSnapshot, filters: QueryFilters) {
} }
} }
/**
* Standup — the morning ritual as a compact, model-narratable payload: today's
* plan (the scheduler's earliest pick per person), overnight drift (real
* anomalies — issues sitting in review or steeping past their estimate), and the
* single stalest blocker to nag about. All from deterministic code; the model
* narrates it. Satisfies #28's "surfaces schedule drift + at least one stale blocker".
*/
function standupView(snap: ProjectSnapshot, asOf: Date) {
const plan = schedule(toSchedulable(snap.issues), snap.deps)
const open = snap.issues.filter((i) => i.state === 'open')
const infOf = (i: GiteaIssue) => inferLifecycle(i, snap.timelines[i.number] ?? [], asOf)
// plan: the earliest scheduled pick per assignee (dependency + priority order)
const seen = new Set<string>()
const planPicks: { who: string; issue: number; title: string; why: string }[] = []
for (const item of plan.items) {
const who = open.find((i) => i.number === item.number)?.assignee ?? 'unassigned'
if (seen.has(who)) continue
seen.add(who)
planPicks.push({
who,
issue: item.number,
title: item.title,
why: item.critical
? 'on the critical path'
: item.blocks.length
? `blocks ${item.blocks.map((b) => `#${b}`).join(', ')}`
: 'next by dependency + priority',
})
}
// drift: real overnight anomalies (review-sitting, steeping past estimate)
const drift: { issue: number; note: string }[] = []
for (const i of open) {
const inf = infOf(i)
if (inf.column === 'review') drift.push({ issue: i.number, note: `#${i.number} is sitting in review` })
else if (inf.steepingDays != null && inf.steepingDays > (i.facts.estimateDays ?? 2))
drift.push({
issue: i.number,
note: `#${i.number} has steeped ${inf.steepingDays}d past its ${i.facts.estimateDays ?? 2}d estimate`,
})
}
// nag: the single longest-steeping open issue + what it blocks
let nag: { issue: number; steepingDays: number; blocks: number[] } | null = null
for (const i of open) {
const inf = infOf(i)
if (inf.steepingDays == null) continue
if (!nag || inf.steepingDays > nag.steepingDays) {
nag = { issue: i.number, steepingDays: inf.steepingDays, blocks: plan.items.find((it) => it.number === i.number)?.blocks ?? [] }
}
}
return { date: asOf.toISOString().slice(0, 10), plan: planPicks.slice(0, 5), drift: drift.slice(0, 5), nag }
}
/** Build the compact payload for one view. Unknown/unbuilt views return a marker. */ /** Build the compact payload for one view. Unknown/unbuilt views return a marker. */
export function buildProjectView( export function buildProjectView(
view: ProjectView, view: ProjectView,
@@ -140,6 +196,8 @@ export function buildProjectView(
return issueView(snap, f, asOf) return issueView(snap, f, asOf)
case 'search': case 'search':
return searchView(snap, f) return searchView(snap, f)
case 'standup':
return standupView(snap, asOf)
default: default:
return { notImplemented: view } return { notImplemented: view }
} }

View File

@@ -111,6 +111,8 @@ export {
} from './agent/agent-tools.js' } from './agent/agent-tools.js'
export { buildProjectView } from './agent/query-project.js' export { buildProjectView } from './agent/query-project.js'
export type { ProjectSnapshot, ProjectView, QueryFilters } from './agent/query-project.js' export type { ProjectSnapshot, ProjectView, QueryFilters } from './agent/query-project.js'
export { activeDirectives, assembleHotContext, estimateTokens, HOT_CONTEXT_BUDGET_TOKENS } from './agent/memory-v0.js'
export type { HotContextInputs } from './agent/memory-v0.js'
export { CAPTURE_SYSTEM, captureWork, parseCaptureArgs, PROPOSE_ISSUES_TOOL } from './agent/capture-work.js' export { CAPTURE_SYSTEM, captureWork, parseCaptureArgs, PROPOSE_ISSUES_TOOL } from './agent/capture-work.js'
export type { CaptureProposal, ProposedIssue } from './agent/capture-work.js' export type { CaptureProposal, ProposedIssue } from './agent/capture-work.js'

View File

@@ -0,0 +1,75 @@
/**
* Performance pass (#32). The deterministic compute path must stay well under the
* PLAN.md targets on representative fixtures:
* - scheduler + Monte Carlo forecast < 1s @ 200 open issues.
* - scaling stays roughly linear (no accidental O(n²) in the hot path).
*
* Reconcile-<5s@500 is network-bound (~2N gitea calls) and is covered by the live
* reconcile, not here — this file benchmarks the pure compute the app runs each
* turn. Bounds are the actual targets with comfortable headroom so timing jitter
* can't flake the suite; actuals are logged.
*/
import { describe, expect, it } from 'vitest'
import { forecast } from '../forecast/forecast-v0.js'
import { type DependencyEdge, schedule, type SchedulableIssue } from '../scheduler/scheduler-v0.js'
import { scheduleWithCapacity, type Worker } from '../scheduler/scheduler-capacity-v0.js'
const EST = [1, 2, 3, 5, 8]
const WORKERS: Worker[] = [
{ person: 'a', speed: 0.8 },
{ person: 'b', speed: 0.6 },
{ person: 'c', speed: 1.0 },
]
/** A representative open backlog: varied estimates/priorities/assignees + a light dependency web. */
function backlog(n: number): { issues: SchedulableIssue[]; edges: DependencyEdge[] } {
const issues: SchedulableIssue[] = Array.from({ length: n }, (_, i) => ({
number: i + 1,
title: `Issue ${i + 1} with a representative title of some length`,
labels: [`est/${EST[i % EST.length]}d`, `p/${(i % 4) + 1}`],
estimateDays: EST[i % EST.length],
priority: (i % 4) + 1,
assignee: WORKERS[i % WORKERS.length].person,
}))
// ~1 dependency per 3 issues, always on a lower-numbered issue (acyclic)
const edges: DependencyEdge[] = []
for (let i = 3; i < n; i += 3) edges.push({ issue: i + 1, dependsOn: i - 1 })
return { issues, edges }
}
function ms(fn: () => void): number {
const t0 = performance.now()
fn()
return performance.now() - t0
}
describe('perf (#32)', () => {
it('scheduler + Monte Carlo forecast < 1s @ 200 open issues', () => {
const { issues, edges } = backlog(200)
const elapsed = ms(() => {
schedule(issues, edges)
scheduleWithCapacity(issues, edges, WORKERS)
forecast(issues, edges, { workers: WORKERS }) // 2000 trials (default)
})
// eslint-disable-next-line no-console
console.log(`[perf] schedule+capacity+forecast @200 = ${elapsed.toFixed(1)}ms`)
expect(elapsed).toBeLessThan(1000)
})
it('scales roughly linearly — 400 issues is well under 4x the 100-issue time', () => {
const small = backlog(100)
const big = backlog(400)
const run = (b: typeof small) => () => {
schedule(b.issues, b.edges)
forecast(b.issues, b.edges, { workers: WORKERS })
}
// warm up (JIT) so the ratio reflects steady state
run(small)()
const t100 = Math.max(ms(run(small)), 0.1)
const t400 = ms(run(big))
// eslint-disable-next-line no-console
console.log(`[perf] @100 = ${t100.toFixed(1)}ms · @400 = ${t400.toFixed(1)}ms · ratio ${(t400 / t100).toFixed(1)}x`)
expect(t400).toBeLessThan(t100 * 8) // generous: rules out O(n²), tolerant of jitter
})
})