- Smoke tests call non-existent CLI commands → test the underlying function directly via
tsx - Example: instead of
npx tsx src/index.ts route, importclassifyTaskfromsrc/ai/task-classifier.tsand call it - Check
src/cli.tsto confirm a CLI command actually exists before using it
- Import functions directly:
import { X } from './src/path.js' - Parse return values, not stdout strings
- This is faster (no CLI startup overhead) and more precise
- Node module path:
node ./node_modules/tsx/dist/cli.mjs <file.ts> - Works for both
.tsand.mjsfiles - Use temp files for inline tests, clean up in
finallyblock
- Counting function occurrences in source files breaks when code is refactored
- Instead, test actual behavior: call the function and check the return value
- Source-level checks are OK for "does file X exist" but not for "how many times is Y referenced in file Z"
Start-Processcreates child processes that get killed when parent PowerShell exits- Use
cmd /c start "" "program.exe" argsto fully detach - This is important for 9Router, web servers, and any process started by a smoke
- The
tests/directory has 40+ test files covering routing, chat, verification, etc. - Smokes should delegate to existing unit tests via
run-tests.ps1or directtsx --testcalls - Don't duplicate test logic in smoke scripts
- Smokes support
-Quickto skip slow direct function tests - Run only the essential checks in quick mode; defer full tests to normal mode
Each smoke should:
- Print
"=== smoke:<name> ==="header - For each check:
[OK]or[FAIL]prefix, colored output - Print
"=== smoke:<name> PASSED/FAILED ==="at end - Exit with 0 for pass, 1 for fail
- Generated code/task answers should never tell the user to manually run
npm test,npm run build,npx tsc, or smoke scripts - If OpenCode has terminal access, it must run verification steps itself
- The answer quality critic (
src/ai/answer-quality.ts) flagsmanual_verification_requestfor any prompt suggesting manual test execution
- Any generated answer suggesting
delete,remove,rm,format,shutdown, or similar destructive actions must also include approval-seeking wording (shall I,do you want me to,approve,can I) - The quality critic flags
unsafe_action_without_approvalwhen destructive commands lack approval phrasing
- If a smoke tests individual functions/helpers instead of a real E2E flow, its name and documentation must reflect this
smoke:hysa-chatis component-level (tests provider routing helpers + unit tests), not true E2E chatsmoke:hysa-chat-e2eis the true full-stack chat check (starts web server, POSTs to real chat endpoints). Default mode is deterministic (usesHYSA_E2E_TEST_PROVIDER=true).- True E2E smokes must be named with
-e2esuffix
write_file,run_command, and any tool with risk levelreviewordangerousrequire explicit approval- Tools must check
ToolRunContext.approvedbefore executing destructive actions run_commandmust useisDangerousCommand()to block destructive patterns (rm -rf, format, shutdown, etc.) even with approval- The
--approveflag in CLI must never override dangerous command blocking
write_fileandrun_commandmust return proposed action details in dry-run mode without executingwrite_filedry-run must include diff output if file existsrun_commanddry-run must show the command string, safety classification, and cwd- Dry-run should be the default; explicit
approved=trueis required for actual execution
- Action log entries must redact: API keys, tokens, passwords, private keys, credentials
- Input summaries must be truncated to 500 chars max before logging
- Full file contents must never be logged as input summary
- Logging failures must be non-fatal — tool execution continues with warning
- All file path tools must resolve and validate paths against a project root or cwd
isWithinCwd()must be called before any file operation- Path traversal attempts must return a structured error, not throw
- Binary files must be detected and blocked from
read_file
planToolActionsForTask()must be called before any tool execution in an agent loop- The plan must be reviewable and human-readable via
formatPlanForDisplay() - Blocked actions must never be executed
- In both tool system and agent planner, write_file and run_command must always require approval
- Their
approvalPolicymust never be'auto'in any context statusmust always be'requires_approval'in all plans
- Tool plans must be deterministic (no AI calls during planning)
- Task classification must be pattern-based, not AI-based
- Plans must include per-action risk level, approval policy, status, and reason
- The
--jsonflag must produce deterministic output for testing
- All future E2E smoke tests that start the HYSA web server must use
scripts/lib/hysa-test-server.ps1 - Never use
cmd /c startwith detached processes in smoke scripts - Never duplicate server startup logic inline — dot-source the shared harness instead
- The shared harness provides:
- Managed process (not detached) with PID tracking
- stdout/stderr log capture to
%TEMP%\hysa-test-server\ - Readiness polling on
GET /api/statuswith timeout Write-HysaServerDiagnosticsfor failure diagnosticsStop-HysaTestServerinfinallyblock for cleanup
- Every smoke script that starts the HYSA web server must call
Write-HysaServerDiagnosticson startup failure - Must print: command, cwd, port, readiness URL, process ID, exit code, last stdout/stderr lines
- Use the shared harness functions instead of ad-hoc diagnostics
- Smokes that make live provider calls (9Router probes, Arabic chat) must support
-Quickflag - Quick mode: run essential deterministic checks only, skip slow provider-dependent probes
- Full mode: run all checks including provider probes
- Package scripts must provide both:
smoke:<name>(full) andsmoke:<name>:quick(quick) - CI should use
:quickvariants for provider-dependent smokes
- No
while ($true)loops without a max iteration count or time budget - Every
Invoke-RestMethod/Invoke-WebRequestmust include-TimeoutSec - Every
tsx --testcall must have a timeout wrapper (viascripts/lib/tsx-runner.ps1) - Every server readiness poll must have a configurable timeout
- On timeout, the error must report: which step timed out, what was being waited for, how long waited
smoke:hysa-chat-e2e— Default: deterministic. REQUIRED for CI gate. UsesHYSA_E2E_TEST_PROVIDER=true. Tests real/api/chat+/api/chat/streampaths, sessionId, answerQuality, Arabic routing. Passes reliably every time with zero external provider calls.smoke:hysa-chat-e2e:deterministic— Same as default (alias).smoke:hysa-chat-e2e:live— Live-provider mode. English + Arabic + streaming. Rate-limited responses show[RATE_LIMIT](yellow warning), not[FAIL].smoke:hysa-chat-e2e:live:quick— Live-provider quick mode. English chat only, skips Arabic + streaming.smoke:hysa-chat-e2e:live:required— Live-provider mode that treats rate_limit as[FAIL]. Use this to require a working live provider.- Deterministic E2E is the REQUIRED CI gate. Live provider smokes are provider-dependent.
- When all probe models return
429 Too Many Requests, report clearly: "All N models rate-limited" - Do not hang waiting for rate limits to expire
- In Quick mode, probe at most 3 models
- Always report: URL, model count, models tried, status per model, total duration
HYSA_E2E_TEST_PROVIDER=trueactivatessrc/ai/test-client.tswhich returns deterministic responses ("OK", Arabic "حسنًا")- The test client intercepts in
createClient()BEFORE any provider routing (smart router, fallback chain) - Never activates in normal production — only when the env var is explicitly set
- No API keys, no network calls, no external dependencies
- The test provider is NOT in the
ProviderTypeunion — it cannot be selected by config or smart routing - All
/api/chatpaths (streaming, non-streaming, continueChat, vision fallback) are covered
- Default (no
-TreatRateLimitAsWarning): rate-limited responses are[FAIL]— used bysmoke:hysa-chat-e2e:live:required - With
-TreatRateLimitAsWarning: rate-limited responses are[RATE_LIMIT](yellow warning), not[FAIL]— used bysmoke:hysa-chat-e2e:live - Rate_limit detection: response has a message but no
answerQualityfield AND message contains "rate limit", "busy", "unavailable", "cooldown", "all free", or "timed out"
getMemoryContextForTask({ task, taskKind? })returnsMemoryContextResultwith:recentMemories,relevantMemories,projectFacts,summarymemoryUsed: booleanandmemoryHits: numberrelevantFiles: string[]— file paths extracted from memory items
- Deterministic — no AI calls, wraps existing
selectContext()from brain/context-selector - Safe when memory unavailable (returns empty result, never throws)
- Accepts optional
memoryContextparameter - Plan output includes:
memoryUsed,memoryHits,memoryReasoningfields - Memory-implied files replace generic
list_filesfallback when no user-specified files exist - User-specified files always take priority over memory-implied files
memoryReasoningprovides deterministic explanation (e.g., "Memory shows recent work in: src/web/api.ts")- All existing approval guarantees maintained — memory only influences file prioritization
User Request → getMemoryContextForTask() → MemoryContextResult
↓
planToolActionsForTask({ userText, memoryContext }) → AgentToolPlan
↓
memoryUsed, memoryHits, memoryReasoning in plan
↓
Multi-step agent propagates memory metadata to result
tests/memory-context.test.ts— 10 tests covering empty, shape, Arabic, deterministic, error handlingtests/memory-aware-planner.test.ts— 15 tests covering backward compat, metadata, prioritization, deterministic, path traversal filteringscripts/smoke-memory-aware-agent.ps1— 6 checks (unit tests, metadata, prioritization, deterministic, multi-step integration, package)
POST /api/agent/plan-toolsreturns a deterministic plan (pattern-based, no AI calls)POST /api/agent/execute-toolsonly executes actions whose IDs are inapprovedActionIds- The server stores the original plan in a
Map<string, AgentToolPlan>— the client cannot mutate toolName, input, or parameters - Blocked actions (dangerous commands) never execute, even if the client includes their IDs in
approvedActionIds write_fileandrun_commandalways haveapprovalPolicy: 'requires_approval'andstatus: 'requires_approval'in every plan- Plan-tools is called automatically after every AI response, but only displays the plan panel — never auto-executes
- Auto-continue after execution is only triggered when the user clicks "Execute Approved Actions" in the ToolPlanPanel