Files
grant-outreach-engine/apps/outreach-worker
Croissant Le Doux 4f35dcdb8d feat(profiler): Stage 4 grounded research profiler — cited profiles replace NTEE stubs
profileOrgs (nightly 05:45, 25/run, easy-win orgs best-first per the
plan's coarse-match-gates-profiling rule): two-pass Gemini flash —
Google Search-grounded research (groundingMetadata = citation universe)
then JSON-schema extraction citing only from it. Overwrites the stub
profile (mission/programs/geography/funders/staff/news/sources/
confidence); re-embeds mission fit from researched text unless identity
unconfirmed or confidence <0.5 (then the NTEE embedding stays —
unverified research must not steer scoring). Detail page renders
researched programs w/ source links, funders, staff.

Fixes: org_profiles.confidence is float4 — 0.2 stores as 0.20000000298
so 'confidence <= 0.2' excluded every stub (epsilon comparison); RR7
loader serialization turns unknown into never (hoisted casts).

Live: 55 orgs profiled (41 research-grade, 2 low-confidence, 0 failed).
Scoring effect verified both directions: Annie's Angels x NIH mammalian
models fit 19->12; Seacoast Pathways x Foundation for Seacoast Health
fit 25/30 — and the researched knownFunders independently names that
funder. 161 tests green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 00:08:12 -04:00
..

@novelpad/outreach-worker

DBOS executor process for the HelmDocs grant-match outreach engine. Runs the scheduled workflows that keep the outreach schema's grant catalog fresh — nightly ingestion from upstream sources and an hourly expiry sweep — decoupled from any HTTP-facing app in this monorepo (mirrors the workflow-worker split in novelpad-desktop).

Boot order (load-bearing)

src/main.ts imports ./workflows/ingest-grants.js and ./workflows/expire-grants.js before calling DBOS.launch(). Each of those modules calls DBOS.registerWorkflow + DBOS.registerScheduled at module-evaluation time — DBOS only dispatches scheduled/queued jobs for functions that were registered before launch, so importing them after launch (or not at all) silently means the cron jobs never fire.

The db handle is then threaded into each workflow module via its set*Deps injector (setIngestGrantsDeps / setExpireGrantsDeps), also before launch — DBOS serializes scheduled-function arguments, so a Drizzle client can't be passed through the scheduler call itself. This is the same module-scope-registry pattern novelpad-desktop's workflow-worker uses for setStartDeps.

1. import workflow modules       → registers ingestGrants / expireGrants
2. build pg Pool + drizzle(db)
3. set*Deps({ db })              → populates each workflow's registry
4. DBOS.setConfig(...)
5. DBOS.launch()                 → scheduler starts firing

Workflows

  • ingestGrants (0 3 * * *, nightly) — fetchGrantsGov (stub; real implementation calls the Grants.gov Search2 API, POST https://api.grants.gov/v1/api/search2) → normalize → upsert via serverInsertGrants from @novelpad/outreach-core/server. NH state postings and 990-PF extracts land as additional fetch+normalize steps later.
  • expireGrants (0 * * * *, hourly) — marks grants whose close date has passed as closed via serverExpireClosedGrants, so they drop out of the active match/scoring pool.

Both are registered as a DBOS workflow and a scheduled function referencing the same function object (DBOS.registerWorkflow then DBOS.registerScheduled) — see the doc comments in src/workflows/*.ts for why the dual registration is required.

Scripts

  • yarn devnode --env-file=.env --env-file-if-exists=.env.local --import tsx/esm ./src/main.ts
  • yarn build — esbuild bundle to build/main.js (--packages=external)
  • yarn start — run the built bundle
  • yarn typechecktsc --noEmit

Environment

Copy .env.example to .env and fill in DATABASE_URL (required — main.ts throws on boot without it) and GCP_SERVICE_ACCOUNT_KEY_PATH (used by future ingestion/scoring steps that call Google-hosted APIs).