Skip to content

feat(evaluate): add TypeScript HR assistant agent evaluation sample - #2165

Open
BharathiSrini wants to merge 1 commit into
awslabs:mainfrom
BharathiSrini:typescript_agents
Open

BharathiSrini wants to merge 1 commit into
awslabs:mainfrom
BharathiSrini:typescript_agents

Conversation

@BharathiSrini

Copy link
Copy Markdown
Collaborator

Introduces a complete evaluation workflow for a TypeScript-based HR assistant agent running in AgentCore Runtime:

  • hr-assistant/: Claude Bedrock Agent in TypeScript with OpenTelemetry instrumentation (Dockerfile, package.json, tsconfig.json, src/agent.ts)
  • deploy.py: build, push to ECR, create AgentCore runtime
  • evaluate.py: on-demand evaluation with built-in + code-based evaluators
  • cleanup.py: tear down all provisioned resources
  • .gitignore: allow Dockerfile for this sample

Amazon Bedrock AgentCore Samples Pull Request

Important

  1. We strictly follow a issue-first approach, please first open an issue relating to this Pull Request.
  2. Once this Pull Request is ready for review please attach review ready label to it. Only PRs with review ready will be reviewed.

Issue number:

Concise description of the PR

Evaluations on agents that are developed using TypeScript

Changes to ..., because ...

User experience

Please share what the user experience looks like before and after this change

Checklist

If your change doesn't seem to apply, please leave them unchecked.

  • [ x] I have reviewed the contributing guidelines
  • [ x] Add your name to CONTRIBUTORS.md
  • [ x] Have you checked to ensure there aren't other open Pull Requests for the same update/change?
  • Are you uploading a dataset?
  • [x ] Have you documented Introduction, Architecture Diagram, Prerequisites, Usage, Sample Prompts, and Clean Up steps in your example README?
  • [x ] I agree to resolve any issues created for this example in the future.
  • [x ] I have performed a self-review of this change
  • [ x] Changes have been tested
  • [ x] Changes are documented

Acknowledgment

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.

mvangara10
mvangara10 previously approved these changes Oct 7, 2026
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Latest scan for commit: a525a3d | Updated: 2026-10-09 20:50:06 UTC

Security Scan Results

@BharathiSrini
BharathiSrini force-pushed the typescript_agents branch 2 times, most recently from d54599e to 3893e90 Compare October 9, 2026 19:45
Demonstrates the full AgentCore Evaluations lifecycle for a LangGraph
TypeScript agent deployed as a custom container on AgentCore Runtime:
on-demand (EvaluationClient), dataset (OnDemandEvaluationDatasetRunner),
batch (StartBatchEvaluation) and online (CreateOnlineEvaluationConfig).

- agent.ts: stamp ADOT-style resource attributes (service.name,
  cloud.resource_id, ...) on exported span docs so batch and online
  evaluation can discover sessions; strip Nova <thinking> tags; remove
  unused helpers and dependencies
- deploy.py: inject OTEL_LOG_GROUP_NAME, OTEL_SERVICE_NAME and
  AGENT_RUNTIME_ARN after runtime creation; refactor into functions
- evaluate.py: restructure into per-stage functions with --modes and
  --session-id; add 4-scenario ground-truth dataset; scope batch jobs
  to the run's session IDs; move HRPolicyAccuracy to the dataset stage
  (plain-string expected_response only covers the last trace)
- cleanup.py: delete batch evaluations, online configs and their
  results log groups, evaluators, the runtime and its log group, ECR
  repository and IAM roles recorded in results/cleanup_state.json;
  safe to re-run
- READMEs: document the lifecycle, per-stage evaluators and results
- Pass ruff check/format and pylint (10.00/10)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants