Help

Eval harness guide

Team-facing documentation for authoring, running, reviewing, and comparing evals.

Eval Harness Guide

This harness is the team workspace for creating eval cases, running repeatable experiments, reviewing failures, and comparing improvements over time.

For the folder-level map of CLI entrypoints, migrations, seeds, and local DB files, see evals/README.md.

The browser UI lives in apps/evals and normally runs at http://localhost:3000/evals. The standalone chat frontend lives in apps/chat; its local /api/chat proxy points at the evals runtime through PXCHAT_EVALS_BASE_URL and can pin a registered agent setup with PXCHAT_AGENT_SETUP_ID. Production chat and evals processes must share PXCHAT_PROXY_SECRET for that server-side pin. Browser-supplied setup/model overrides are stripped, and the evals API rejects unauthenticated overrides unless a trusted local instance explicitly enables PXCHAT_ALLOW_RUNTIME_OVERRIDES=1.

What It Is For

  • Writing new eval cases directly in this eval UI.
  • Organizing cases into reusable presets for smoke, regression, critical, or scenario runs.
  • Launching a run only when the environment is actually ready.
  • Reviewing only the cases that need manual attention.
  • Comparing a new run against a previous run or a pinned baseline.

Quick Glossary

  • Case
    • One eval entry. Usually this is one question or one prompt plus the expected answer and metadata around it.
  • Scenario
    • A multi-turn conversation made of several related cases. A scenario key groups those turns together in order.
  • Source
    • The external origin of the case, such as a spreadsheet row, QA artifact, or import identifier. This is for traceability, not runnability.
  • workflowStatus
    • The harness status the team controls directly.
    • draft means still being written.
    • active means runnable.
    • disabled means kept for reference but excluded from runs.
  • sourceStatus
    • A status preserved from the original import workflow or spreadsheet. This stays visible for QA context, but it does not decide whether the harness runs the case.
  • Preset
    • A saved run pack. It stores either a case list or a filter set plus default run options.
  • Baseline
    • The pinned run the team uses as the reference point for comparison.

Main Screens

The eval area now has a shared shell. The top navigation exposes Overview, Cases, Run Builder, Presets, Runs, Setup, and Help, plus quick links back to the configured chat frontend and to the Overview preset launcher. The Setup menu groups Criteria, Agents, Signals, Data Sources, and Models.

  • /evals
    • Overview dashboard.
    • Shows readiness, latest run summary, suite coverage, and a quick launcher for saved preset runs.
  • /evals/runs
    • Runs list.
    • Use quick filters for search, suite, and status; open advanced filters for mode, language, scorer, date range, and result limit.
  • /evals/suites/[suiteSlug]/cases
    • Cases view and scenario authoring.
    • Use this to browse, filter, duplicate, and create cases.
  • /evals/suites/[suiteSlug]/cases/new
    • New-case form.
    • Use this when you want to add a test without importing a file.
  • /evals/suites/[suiteSlug]/runs/new
    • Run Builder.
    • Use this to compose custom runs from filters or explicit cases, validate readiness, choose agent/model variants, and launch.
  • /evals/presets
    • Preset manager.
    • Use this to view, create, update, and delete suite presets, and to attach recurring schedules to them.
  • /evals/criteria
    • Judge criteria manager.
    • Use this to view the five scoring dimensions, edit definitions/examples, create or duplicate criteria sets, and choose the default rubric for new runs.
  • /evals/agents
    • Agent setup manager.
    • Use this to create versioned agent setups, tune the top-level Pi prompt and flat tool registry, select the signals enabled for each setup, expand the instance signal table, and choose the default setup used by the evals runtime /api/chat.
  • /evals/signals
    • Backend-instance signal registry.
    • Use this to create, describe, edit, or deactivate the allowlisted signals available to agent setups on this eval backend.
  • /evals/data-sources
    • Mirage source manager.
    • Browse and preview the active /sources tree, select a destination folder, and atomically import an LLM-wiki ZIP as a new immutable revision.
  • /internals/models
    • Model profile and default model configuration.
    • Linked from the eval shell in the Setup menu as Models.
  • /evals/runs/[runId]
    • Live run page.
    • Shows queued/running state, exception review queue, comparisons, and advanced diagnostics.
  • /evals/help
    • This guide, rendered inside the eval shell.

Agent Runtime And Tools

The evals runtime now uses one top-level Pi agent for both public chat and headless eval execution. An agent setup has one system prompt and one flat tools registry; there is no separate AI SDK main-agent layer or nested Pi tool configuration.

Each tool can be enabled independently:

  • ragSearch queries the embedded Postgres-backed knowledge base.
  • pricingTable2024 queries the FY2024/2025 electricity pricing table.
  • signal sends an allowlisted frontend signal from this backend instance's registry. A setup must enable both the tool and each individual signal it may send.
  • read, grep, find, ls, and bash inspect the read-only Mirage /sources filesystem. Prefer the dedicated file tools; enable bash only when shell pipelines are useful.

The signal registry starts empty for each eval SQLite backend; migrations never seed product-specific signals. Manage the registry from the expandable table on /evals/agents or from /evals/signals. Adding a signal makes it available for setup selection but does not enable it anywhere. The signal tool and every setup's selected-signal list are empty or disabled by default. System prompts and tool schemas include only the active signals selected by that setup.

The /sources filesystem is the active revision owned by this backend's PXCHAT_DATA_SOURCES_DIR; it is not the repository checkout. Chat requests pin a revision for their Pi session, and eval runs pin one revision across every case and scenario turn. ZIP uploads affect Mirage only: ragSearch remains a separate Postgres-backed store. Use a distinct PXCHAT_DATA_SOURCES_DIR, EVAL_DB_PATH, and, when RAG is enabled, DATABASE_URL for each backend use case.

Imports accept UTF-8 text files and reject unsafe paths, symlinks, binary content, name/type collisions, empty archives, and oversized or suspiciously compressed archives. Immutable revisions hard-link unchanged files. The backend retains three inactive revisions by default and refuses additional imports rather than deleting a revision that a session may have pinned; configure PXCHAT_DATA_SOURCES_MAX_INACTIVE_REVISIONS (0–100) when a longer operator-managed history is needed.

Every Data Sources API request requires PXCHAT_DATA_SOURCES_ADMIN_TOKEN; the UI keeps an entered token in sessionStorage only. For local development, PXCHAT_DATA_SOURCES_ALLOW_UNAUTHENTICATED_DEV=1 can explicitly opt out, but it is ignored in production. Writes also default off, so enable PXCHAT_SOURCES_WRITE_ENABLED=1 only for trusted admins. Protect the rest of the eval/admin app separately because it does not provide built-in user authentication.

A setup with every tool disabled is valid. Pi still answers as an ordinary conversational assistant using the selected model; zero tools does not switch the UI into a coding-agent experience or produce a setup warning. The public chat stream exposes the validated final assistant text plus explicitly enabled transient data-signal parts; it does not surface hidden reasoning, intermediate turns, raw tool events, internal /sources paths, or source footers. Headless eval artifacts retain citations and emitted signals for validation and diagnostics.

The chat frontend handles data-signal outside conversation history. Web code can register a handler or listen for the pxchat:signal custom event. An embedded mobile host receives the same { type: "pxchat.signal", signal } envelope through iOS window.webkit.messageHandlers.pxchatSignal.postMessage(...) or Android window.PXChatSignalBridge.postMessage(...). Hosts should allowlist signal names and deduplicate by signal.id; payloads remain untrusted agent output. The built-in web reaction map is intentionally empty until product behavior for a registered signal is approved.

Headless eval runs use the same runtime and configuration, but their per-case artifacts retain internal trace, tool-call, and tool-result telemetry for diagnostics and scoring. This telemetry is available to eval review surfaces without being sent to the end user's chat transcript.

Case Model

Each case now has two status concepts:

  • workflowStatus
    • draft: still being authored, not runnable.
    • active: included in runnable selections.
    • disabled: kept in the library but excluded from runs.
  • sourceStatus
    • Preserved from imports or spreadsheet workflows.
    • Useful as metadata and filtering, but not used to decide if a case should run.

Other important fields:

  • caseType: golden, scenario_turn, or exploratory
  • sourceCaseId: external/spreadsheet identifier
  • scenarioKey + scenarioTurn: used for multi-turn continuity tests
  • isReleaseCritical: marks cases that should weigh heavily in release confidence
  • tags: reusable labels for viewpoints, policy buckets, regression packs, or any other run selector
  • expectedSignals: optional exact signal assertion. null means signals are not graded, [] means no signal is expected, and a list requires exactly those names once each. Missing, unexpected, or duplicate signals fail the case even when answer judging passes.

Run manifests snapshot the active signal definitions. The case form can opt into exact signal assertions without changing legacy cases. JSON imports accept nullable expectedSignals, and Excel imports use optional Expected_Signals. Omitting the field preserves legacy no-grading behavior.

Creating Cases In This Eval UI

  1. Open the case library.
  2. Click New case.
  3. Fill in:
    • caseKey
    • caseType
    • workflowStatus
    • prompt and expected answer fields
  4. Save as draft while refining.
  5. Move to active once the case is ready to run.

Tips:

  • Use Duplicate on an existing case to create a variant quickly.
  • For scenario tests, define a scenarioKey and scenarioTurn.
  • On a scenario-turn case, use Add next turn to branch the next turn from the current one.

Scenario Authoring

The case library includes a scenario-group form.

Use it to define:

  • scenarioKey
  • display names
  • shared targets
  • shared tags

Then add scenario-turn cases that reference that scenarioKey.

Presets

Presets are reusable saved run selections.

They can store:

  • explicit case IDs
  • filter-based selections
  • default mode
  • default language
  • default sample count
  • optional judge criteria set

The overview dashboard supports:

  • choosing a suite and saved preset
  • reviewing the saved mode, language, sample count, variants, and selection style
  • running the preset directly after readiness validation
  • jumping to the Run Builder when a custom run is needed

The dedicated presets page supports:

  • viewing all presets in one place
  • editing preset defaults and selection logic
  • switching a preset between filter-based and explicit-case selection
  • scheduling recurring runs for a preset on this server

Fresh suites include three top-level Pi starter presets: RAG only, Sources only, and RAG + Sources.

Judge Criteria

The criteria page manages the rubric used by the automated judge.

The default criteria set is created automatically by eval DB bootstrap and covers:

  • accuracy
  • relevance
  • consistency
  • safety
  • style

Each dimension stores a definition, one-to-five score anchors, and examples. Criteria sets also store pass thresholds such as minimum total score, minimum safety, minimum accuracy, minimum relevance, and strict language/format gates.

New runs use the default criteria set unless the Run Builder selects another one. Run manifests snapshot the selected criteria so historical runs keep the rubric they used even if the default changes later.

Scheduled Preset Runs

Preset schedules live in the presets page.

Each schedule belongs to one preset and stores:

  • cadence (daily, weekdays, or weekly)
  • local time
  • timezone
  • enabled or paused state
  • last run
  • next scheduled run

When the Next.js service is running, the eval scheduler checks for due schedules and launches the preset automatically.

Scheduled runs still appear in the normal Runs list and run detail pages, so the team can review them the same way as manually launched runs.

Readiness Checks

The harness automatically checks:

  • chat model connectivity
  • judge model connectivity
  • eval SQLite DB access
  • app Postgres access
  • embedding/preflight coverage

Launch behavior:

  • blocking failures stop a run from starting
  • coverage issues remain warnings unless you later decide to enforce them more strictly

The SQLite eval DB is local machine state by default. The repo ignores evals/*.sqlite*, including -wal and -shm sidecars and recovery backups, so a fresh clone gets its own run history. Use bun run eval:db-check before a demo, and use bun run eval:db-checkpoint before copying or backing up a local DB file.

Launching Runs

Use the Overview dashboard when you want to run a saved preset pack quickly.

Use the Run Builder to:

  1. Select explicit cases or define a filtered pack.
  2. Choose agent variants and candidate model variants.
  3. Choose mode, language, sample count, judge model, and judge criteria.
  4. Confirm the readiness summary is clear for blocking checks.
  5. Launch the run.

After launch, the app redirects immediately to the live run page so the team can see that the run actually started.

Reviewing Results

The run detail page is organized into:

  • live execution header
  • release gate and quality metrics
  • comparisons against baseline and previous run
  • exception review queue
  • failed-case summary
  • per-case drilldown
  • advanced diagnostics

The exception queue is the main manual-review surface.

Cases land there when they are:

  • failed
  • errored
  • ungraded
  • judge/gate conflicts
  • manually reviewed

Baselines And Comparison

Use Use as suite baseline on a run page to pin the suite’s reference run.

Each run can then be compared against:

  • the pinned baseline
  • the immediately previous run

The comparison highlights:

  • pass-rate delta
  • score delta
  • error delta
  • changed verdicts
  • changed cases

CLI Commands

These still work and remain useful for imports or headless runs.

bun run eval:migrate
bun run eval:db-check
bun run eval:db-checkpoint
bun run eval:preflight
bun run eval:judge-check
bun run eval:import -- --suite pxchat --file ./evals/seed/your_dataset.xlsx
bun run eval:run -- --suite pxchat --mode isolated --lang ja --samples 1

Recommended Team Workflow

  1. Add or duplicate cases in this eval UI.
  2. Keep new work in draft until the prompts and expectations are ready.
  3. Move stable cases to active.
  4. Save presets for the packs the team will rerun often.
  5. Run the preset.
  6. Review only the exception queue.
  7. Compare the new run against baseline.
  8. Pin a new baseline only when the results are genuinely better.

Troubleshooting

  • Run blocked before launch:
    • Read the readiness summary and advanced details first.
    • Fix chat model, judge, or DB connectivity before retrying.
  • Coverage warning:
    • The run can still start, but retrieval quality may be misleading.
  • Cases not appearing in a run:
    • Check workflowStatus; only active cases are runnable.
  • Imported spreadsheet status looks odd:
    • That value is preserved as sourceStatus; it does not control runnability.