Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Tasks over Application Manuals (TAM)

This is the executable artifact for Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models, accepted at CIKM 2026. It contains the eight baseline implementations, the approved 200-case federal-sentencing benchmark, public manuals, a restricted MIMIC cohort builder, input validation, and the reported aggregate results.

Artifact Contents

The complete 200-case legal benchmark, all public manuals, and the paper's sampling settings are included. MIMIC-IV and MIMIC-IV-Note cannot be redistributed; credentialed PhysioNet users can reconstruct the 38,332-case eligible ICD cohort with the included script.

The ICD experiments use a random 1,000-case sample with seed 42. The standalone evaluators use those settings by default. The reported paper results are preserved in results/paper_results.csv.

Dataset: The complete benchmark is also available on HuggingFace: manulife/tam-benchmarks

Reported Results

Baseline ICD EM ICD Precision ICD Recall ICD Primary Dx Legal EM Legal MAE
Single-pass RAG 1.0% 52.0% 33.0% 42.7% 7.5% 2.87
Agentic RAG 1.0% 59.0% 37.0% 45.7% 11.0% 2.80
ReAct-style tool use 1.0% 58.88% 26.82% 20.0% 15.5% 3.04
Agent harness 0.7% 59.9% 27.4% 50.1% 14.5% 2.34

Reproduction Guide

Complete the following steps in order.

1. Install

Requirements:

  • Python 3.12
  • an OpenAI or Azure OpenAI API key for model-backed evaluation
  • an embedding deployment for the RAG and agentic-RAG baselines
  • authorized MIMIC-IV and MIMIC-IV-Note access to reconstruct the ICD cases

Prebuilt vector indices are not included in this repository. In Step 5, you will generate them from the bundled manuals and store them locally under .tam_indices/; this default setup does not require a separate search service. Azure AI Search remains available as an optional backend. The ReAct-style and agent-harness baselines read the bundled PDF, XML, and HTML manuals directly; they do not require Document Intelligence exports or a search index.

Windows PowerShell:

py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .
Copy-Item .env.example .env

macOS or Linux:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
cp .env.example .env

After activation, the remaining python commands are identical on all three platforms.

2. Configure OpenAI Or Azure

Edit the .env created in Step 1 and configure one of the following provider paths. The RAG index builder and evaluators use the same embedding setting.

Option A: OpenAI

The base installation supports OpenAI directly:

OPENAI_CHAT_MODEL=openai:gpt-5
OPENAI_API_KEY=YOUR_KEY
TAM_EMBEDDING_MODEL=openai:text-embedding-3-large
TAM_SEARCH_BACKEND=local

Option B: Azure OpenAI

Install the Azure integrations:

python -m pip install -e ".[azure]"

Configure Azure OpenAI while retaining the local LangChain vector store:

OPENAI_CHAT_MODEL=azure_openai:YOUR_CHAT_DEPLOYMENT
TAM_EMBEDDING_MODEL=azure_openai:text-embedding-3-large
TAM_SEARCH_BACKEND=local

AZURE_OPENAI_ENDPOINT=https://YOUR_RESOURCE.cognitiveservices.azure.com/
OPENAI_API_VERSION=YOUR_SUPPORTED_API_VERSION
AZURE_OPENAI_API_VERSION=YOUR_SUPPORTED_API_VERSION
AZURE_OPENAI_DEPLOYMENT=YOUR_CHAT_DEPLOYMENT
AZURE_OPENAI_EMBEDDING_DEPLOYMENT=text-embedding-3-large
AZURE_OPENAI_API_KEY=YOUR_KEY

# Forces both Deep Agents evaluators to use the API-key route instead of
# inferring Azure AD from unrelated AZURE_CLIENT_* variables.
ICD_DEEPAGENTS_MODEL_PROVIDER=azure_key

To use Azure AI Search instead of the local vector store, change TAM_SEARCH_BACKEND to azure and add:

AZURE_SEARCH_ENDPOINT=https://YOUR_SEARCH_SERVICE.search.windows.net
AZURE_SEARCH_API_KEY=YOUR_SEARCH_KEY
ICD_INDEX_NAME=icd-rag-baseline
LEGAL_INDEX_NAME=legal-rag-baseline

Use only one provider configuration at a time. Existing shell variables take precedence over .env; start a fresh terminal or clear stale OPENAI_*, AZURE_OPENAI_*, and AZURE_CLIENT_* values before running.

Verify the installation without making a model call:

python scripts/reproduce/run_tam_baseline.py --experiment icd-single-pass-rag -- --help

3. Prepare Data

3.1 Legal Data

The complete 200-case legal benchmark is included at data/final-approved-200/federal_sentencing_legal_final_dataset_approved.csv and is ready to use.

3.2 ICD Data

MIMIC-IV and MIMIC-IV-Note are credentialed PhysioNet datasets. Access is free for approved researchers but must be requested individually.

3.2.1 Obtain PhysioNet Access
  1. Create a PhysioNet account and complete the credentialing application.
  2. Complete the CITI Program Data or Specimens Only Research course. PhysioNet's CITI instructions explain how external researchers can affiliate with Massachusetts Institute of Technology Affiliates for this training.
  3. Download the full CITI completion report from the CITI Records page. Upload the report, not the completion certificate, through PhysioNet training.
  4. After credentialing and training are approved, open the MIMIC-IV project page and accept its current Data Use Agreement.
  5. Open the MIMIC-IV-Note project page and accept its current Data Use Agreement as well.

PhysioNet displays the applicable license, required training, and DUA in the Access and Files sections of each project page. Follow the current terms shown there and do not share credentials or downloaded data.

3.2.2 Download The Required Files

Once both projects show approved access, download these four files from their Files sections and arrange them as follows:

MIMIC_ROOT/
|-- hosp/
|   |-- admissions.csv.gz
|   |-- diagnoses_icd.csv.gz
|   `-- patients.csv.gz
`-- note/
		`-- discharge.csv.gz

Use files from the MIMIC-IV and MIMIC-IV-Note releases you were approved to access, and record those release numbers in the build command.

3.2.3 Build The ICD Cohort

Build the strict cohort with the streaming/SQLite script.

Windows PowerShell:

python scripts/mimic/build_mimic_icd10_dataset_2017_2019.py `
	--admissions-path "C:\path\to\mimic\hosp\admissions.csv.gz" `
	--diagnoses-path "C:\path\to\mimic\hosp\diagnoses_icd.csv.gz" `
	--patients-path "C:\path\to\mimic\hosp\patients.csv.gz" `
	--notes-path "C:\path\to\mimic\note\discharge.csv.gz" `
	--mimic-iv-release "YOUR_MIMIC_IV_VERSION" `
	--mimic-iv-note-release "YOUR_MIMIC_IV_NOTE_VERSION"

macOS or Linux:

python scripts/mimic/build_mimic_icd10_dataset_2017_2019.py \
	--admissions-path "/path/to/mimic/hosp/admissions.csv.gz" \
	--diagnoses-path "/path/to/mimic/hosp/diagnoses_icd.csv.gz" \
	--patients-path "/path/to/mimic/hosp/patients.csv.gz" \
	--notes-path "/path/to/mimic/note/discharge.csv.gz" \
	--mimic-iv-release "YOUR_MIMIC_IV_VERSION" \
	--mimic-iv-note-release "YOUR_MIMIC_IV_NOTE_VERSION"

The default output is data/mimic/mimic_icd10_note_dataset_2017_2019_strict.csv. A matching reconstruction contains 38,332 eligible encounters.

Each ICD evaluator defaults to the paper protocol: random sampling, seed 42, and a 1,000-case limit.

4. Validate Inputs

Validate the included legal files immediately:

python scripts/reproduce/validate_inputs.py

After reconstructing the ICD cohort, validate both domains:

python scripts/reproduce/validate_inputs.py --icd-dataset data/mimic/mimic_icd10_note_dataset_2017_2019_strict.csv

5. Build Vector Indices

Only single-pass RAG and agentic RAG need search indices. With the default local backend, build each persisted index once:

python scripts/run_icd_chunking.py --recreate-index
python scripts/run_legal_chunking.py --years 2021 2022 2023 2024 2025 --recreate-index

These commands parse the bundled manuals, embed each chunk with text-embedding-3-large, and create persistent LangChain vector stores at:

.tam_indices/icd-rag-baseline.json
.tam_indices/icd-rag-baseline.json.meta.json
.tam_indices/legal-rag-baseline.json
.tam_indices/legal-rag-baseline.json.meta.json

Build the indices once. Both RAG baselines in each domain load the corresponding index automatically, so no index-path argument is needed. --recreate-index replaces an existing index; omit it to update or reuse one. A parsing-only check is available with --skip-upload --limit 25 and does not call the embedding provider.

Index construction sends the bundled manual chunks to the configured embedding provider. It can take several minutes and may incur provider charges.

Confirm retrieval before a long evaluation:

python scripts/run_icd_chunking.py --query-text "type 2 diabetes" --top-k 3
python scripts/run_legal_chunking.py --query-text "acceptance of responsibility" --top-k 3

The same controls are available on the four RAG evaluators: --search-backend, --embedding-model, and --local-index-path. The embedding model used for querying must match the model recorded when the index was built.

When TAM_SEARCH_BACKEND=azure, the same build commands create the named indices in your Azure AI Search service. Verify retrieval with:

python scripts/run_icd_chunking.py --search-backend azure --embedding-model azure_openai:text-embedding-3-large --query-text "type 2 diabetes" --top-k 3
python scripts/run_legal_chunking.py --search-backend azure --embedding-model azure_openai:text-embedding-3-large --query-text "acceptance of responsibility" --top-k 3

The paper runs used Azure AI Search hybrid text/vector retrieval. The local LangChain backend uses dense cosine retrieval and provides the simplest reproduction path.

6. Run One Baseline

The wrapper registers local pandas dataframes and forwards arguments after -- to the selected evaluator.

Run one ICD baseline:

python scripts/reproduce/run_tam_baseline.py --experiment icd-single-pass-rag --icd-dataset-path data/mimic/mimic_icd10_note_dataset_2017_2019_strict.csv

Run one legal baseline:

python scripts/reproduce/run_tam_baseline.py --experiment legal-single-pass-rag

Available experiment IDs are:

icd-single-pass-rag
icd-agentic-rag
icd-react-style-tool-use
icd-deepagents
legal-single-pass-rag
legal-agentic-rag
legal-react-style-tool-use
legal-deepagents

Each command runs the baseline's full default evaluation. Agentic RAG, ReAct, and Deep Agents make multiple model calls and can take substantially longer than single-pass RAG.

7. Run All Baselines

The cross-platform runner validates the inputs and then executes all eight baseline/domain combinations in sequence:

python scripts/reproduce/run_all.py --icd-dataset-path data/mimic/mimic_icd10_note_dataset_2017_2019_strict.csv

Use --skip-icd or --skip-legal to run one domain. The script stops on the first failed baseline and returns its exit code. It uses each ICD evaluator's default of 1,000 random cases with seed 42 and the included legal files.

The command uses the model and retrieval settings from .env and runs the full workload, which may take several hours and incur substantial model-provider charges.

8. Inspect MLflow Results

Every evaluator forces standard MLflow tracking to the repository-local mlruns/ directory. Inspect runs with:

mlflow ui --backend-store-uri mlruns

Then open http://127.0.0.1:5000. Legal runs report exact match and mean absolute error; ICD runs report exact set match, per-case precision/recall, and primary-diagnosis accuracy. LLM-backed reruns may differ from the paper because hosted models and service behavior change over time.

Repository Layout

  • baselines/: the eight implementation packages.
  • scripts/evaluate_*.py: direct evaluators and MLflow logging.
  • scripts/reproduce/: the public wrapper, all-eight runner, and input validator.
  • scripts/mimic/: restricted ICD cohort reconstruction.
  • data/final-approved-200/: included legal benchmark.
  • data/reference-manuals/: included public ICD and legal manuals.
  • results/: machine-readable reported results.

Citation

@inproceedings{soni2026tasks,
	author = {Soni, Utkarsh and Murtaza, Syed Shariyar and Nie, Yifan and Chandrasekhar, Sachin and Wen, Eugene},
	title = {Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models},
	booktitle = {Proceedings of the 35th ACM International Conference on Information and Knowledge Management},
	year = {2026},
	address = {Rome, Italy},
	publisher = {Association for Computing Machinery},
	isbn = {979-8-4007-2539-5},
	doi = {10.1145/3799682.3841039}
}

This artifact is released under the terms in LICENSE. Raw MIMIC data and reconstructed patient records are not covered for redistribution.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages