This guide helps triage and debug issues with cmsdev automated tests, particularly focusing on CASMTRIAGE tickets generated from automated test runs.
Auto-triage tickets follow the pattern: CASMTRIAGE-XXXX
Example: CASMTRIAGE-8823
Check the ticket description or title to identify which NCN (Non-Compute Node) the test ran on.
Common node patterns:
ncn-w001- Worker node 1ncn-w002- Worker node 2ncn-m001- Master node 1
- In JIRA ticket: Look for hyperlink with text "Test log output"
- Click the link to open the log viewer page
- Click "stdout" hyperlink to view the actual test logs
- Identify the log directory from the output:
Starting main run, version: 1.34.0, log directory: /opt/cray/tests/install/logs/cmsdev/241012_050305_414367_990773 - Access logs on the node:
ssh ncn-w001 # or appropriate node cd /opt/cray/tests/install/logs/cmsdev/241012_050305_414367_990773 less cmsdev.log
- Check for artifacts (if test failed):
# Artifacts are saved in the same directory ls -lh artifacts.tgz # Extract artifacts tar -xzf artifacts.tgz ls -R artifacts/
- Access the single log file:
ssh ncn-w001 less /opt/cray/tests/install/logs/cmsdev/cmsdev.log
- Find the run tag from JIRA or log output:
Creating temporary directory Starting sub-run, tag: 0P65L-bos - Filter logs by run tag:
# Extract logs for specific run tag grep "run=0P65L" /opt/cray/tests/install/logs/cmsdev/cmsdev.log > /tmp/filtered-0P65L.log # For service-specific sub-run grep "run=0P65L-bos" /opt/cray/tests/install/logs/cmsdev/cmsdev.log > /tmp/bos-test.log
| Aspect | Legacy (< 1.34.0) | Current (>= 1.34.0) |
|---|---|---|
| Log file | Single shared cmsdev.log |
Per-run directory with own cmsdev.log |
| Directory | /opt/cray/tests/install/logs/cmsdev/ |
/opt/cray/tests/install/logs/cmsdev/<timestamp>/ |
| Run isolation | Run tag required to filter | Each run has its own directory |
| Artifacts | In artifacts/ subdirectory per tag |
artifacts.tgz in run directory |
| Timestamp format | N/A | YYMMDD_HHMMSS_microseconds_PID |
Python tests write logs to a separate directory:
/opt/cray/tests/integration/logs/csm/cmstools/
+-- barebones_image_test/
| +-- YYYYMMDD_HHMMSS.log
+-- cfs_sessions_rc_test/
+-- YYYYMMDD_HHMMSS.log
# Find the latest barebones image test log
ls -lt /opt/cray/tests/integration/logs/csm/cmstools/barebones_image_test/ | head -5
# Find the latest CFS race condition test log
ls -lt /opt/cray/tests/integration/logs/csm/cmstools/cfs_sessions_rc_test/ | head -5Log Pattern:
ERROR: Expected 3 or more pods with prefix cray-bos, found 2
or
ERROR: Expected Running/Succeeded phase for pod cray-cfs-api-xxxx, found phase=CrashLoopBackOff
Triage Steps:
kubectl get pods -n services -l app.kubernetes.io/name=cray-bos
kubectl describe pod <failing-pod> -n services
kubectl logs <failing-pod> -n services --previousLog Pattern:
ERROR: Unexpected status code 503 from GET https://api-gw-service-nmn.local/apis/bos/v2/healthz
Triage Steps:
# Check API gateway
curl -k https://api-gw-service-nmn.local/apis/
# Check service health directly
kubectl logs -n services -l app.kubernetes.io/name=cray-bos-apiLog Pattern:
ERROR: Expected Bound status for pvc=cray-console-operator-data-claim, found status=Pending
Triage Steps:
kubectl get pvc -n services <pvc-name>
kubectl describe pvc -n services <pvc-name>
kubectl get pv | grep Available
kubectl get storageclassLog Pattern:
ERROR: Test timeout after 300 seconds
Timeout values per service (from test.go):
- BOS, CFS, conman, iPXE/TFTP, VCS/gitea: 300 seconds
- All other services (default): 120 seconds
Triage Steps:
- Check if retry was enabled:
--retryflag - Review service response times in logs
- Check for network issues or service overload
- Verify Kubernetes cluster health
Log Pattern:
ERROR: Command failed: /usr/bin/cray bos v2 sessiontemplates list --format json
Triage Steps:
# Verify CLI is configured
cray init
# Check CLI configuration
cat /root/.config/cray/configurations/default
# Test CLI manually
cray bos v2 healthz listLog Pattern:
ERROR: Unexpected status code from tenant-scoped request with Cray-Tenant-Name header
Triage Steps:
# Check tenants
kubectl get namespaces | grep tenant
# Verify tenant CRDs
kubectl get tenants -n tenantsWhat the test checks: At least 3 pods (cray-bos prefix), migration pod status (Succeeded), API CRUD (session templates, sessions, components), version/healthz endpoints.
Common Issues:
- Session template creation failures
- Component listing errors
- Migration pod not completing (should be
Succeeded)
Key Logs to Check:
kubectl logs -n services -l app.kubernetes.io/name=cray-bos-api
kubectl logs -n services -l app.kubernetes.io/name=cray-bos-operatorRelated Files: cmsdev/internal/test/bos/
What the test checks: At least 2 pods (cray-cfs prefix), API CRUD (configurations, sessions, sources, components, options), product catalog validation.
Common Issues:
- Configuration CRUD failures
- Product catalog access errors
- Tenant isolation issues
Key Logs to Check:
kubectl logs -n services -l app.kubernetes.io/name=cray-cfs-api
kubectl logs -n services -l app.kubernetes.io/name=cray-cfs-operatorRelated Files: cmsdev/internal/test/cfs/
What the test checks: 1 pod (cray-ims prefix), PVC bound status (cray-ims-data-claim), API CRUD (images, recipes, public-keys), remote image builds (S3 upload, job monitoring).
Common Issues:
- PVC not bound
- Image creation/deletion failures
- S3 upload errors
- Remote build job failures
Key Logs to Check:
kubectl logs -n services -l app.kubernetes.io/name=cray-ims
kubectl get pvc -n services cray-ims-data-claimRelated Files: cmsdev/internal/test/ims/
What the test checks: Pod running, PVC bound (cray-console-operator-data-claim, cray-console-node-0-data-claim), console API access.
Common Issues:
- PVC not bound (frequent)
- Console operator pod not running
- Data claim issues
Key Logs to Check:
kubectl logs -n services -l app.kubernetes.io/name=cray-console-operator
kubectl get pvc -n services | grep consoleRelated Files: cmsdev/internal/test/conman/
What the test checks: At least 2 pods (gitea-vcs prefix), PVC bound status, VCS API (repository CRUD, file operations), authentication via K8s vcs-user-credentials secret.
Common Issues:
- Gitea database issues
- PVC problems
- Repository operations failing
- VCS credentials secret missing or invalid
Key Logs to Check:
kubectl logs -n services -l app.kubernetes.io/name=gitea-vcs
kubectl get secret -n services vcs-user-credentialsRelated Files: cmsdev/internal/test/vcs/
What the test checks: BSS iPXE pods (per-architecture: aarch64 and x86_64), cray-tftp pod, TFTP file upload/download operations.
Common Issues:
- Architecture-specific pods missing
- TFTP transfer failures (only runs on worker NCNs)
Key Logs to Check:
kubectl get pods -n services | grep -E "ipxe|tftp"Note: TFTP file transfer subtests only run on worker NCNs (ncn-w*). Failures on master nodes for this subtest are expected to be skipped.
Related Files: cmsdev/internal/test/ipxe_tftp/
Script location: /opt/cray/tests/integration/csm/barebones_image_test
Test flow: Product catalog lookup -> IMS image creation -> CFS image customization -> BOS session -> Node boot -> Validation -> Cleanup
Common Failures:
| Failure Point | Symptoms | Triage Steps |
|---|---|---|
| Product catalog | "No CSM entry found" | kubectl -n services get cm cray-product-catalog -o jsonpath='{.data.csm}' |
| HSM node lookup | "No enabled compute nodes" | cray hsm state components list --type Node --role Compute --enabled true |
| CFS image customization | Session stuck or failed | Check CFS session status, Ansible logs |
| BOS boot | Session never completes | Check BOS operator logs, node console output |
| Cleanup | Resources left behind | Use --no-cleanup to inspect; manually delete |
Key Arguments: See Barebones Image Boot Test -- Command-Line Arguments in the Test Suites Guide.
Script location: /opt/cray/tests/integration/csm/cfs_sessions_rc_test
Test flow: Scale CFS operator to 0 -> create sessions -> run subtests (concurrent DELETE/GET) -> cleanup -> restore operator
Critical recovery note: If the test is interrupted, the CFS operator may remain scaled to 0:
# Check current replica count
kubectl get deployment -n services cray-cfs-operator
# Manually restore (typically 1 replica)
kubectl scale deployment -n services cray-cfs-operator --replicas=1Subtests: See CFS Sessions Race Condition Test -- Subtests in the Test Suites Guide.
Key Arguments: See CFS Sessions Race Condition Test -- Command-Line Arguments in the Test Suites Guide.
Common Failures:
- CFS operator not scaling back up after test
- CFS v2
page-sizeoption left modified (restore withcray cfs v3 options update --default-page-size 1000) - Unexpected HTTP status codes from concurrent operations (indicates real race condition bugs)
When a cmsdev service test fails, artifacts are collected automatically into artifacts.tgz in the log directory:
The following K8s resources are collected from the services namespace:
- nodes, namespaces, pods, pv, pvc, services
- daemonsets, statefulsets, deployments, etcd
- configmaps, secrets, endpoints, postgresqls
- cronjobs, jobs, sealedsecrets, etcdbackups
For each failed service, detailed information is collected via ArtifactDescribeNamespacePods():
kubectl describe pod -n services <pod> --show-events=true-- for each service podkubectl logs -n services <pod> --all-containers=true --timestamps=true --prefix=true-- for each service podkubectl describe pvc -n services <pvc> --show-events=true-- for each service PVC
RPM query output for: craycli, docs-csm, csm-testing, goss-servers
# Gather all pods status
kubectl get pods --all-namespaces -o wide > /tmp/all-pods.txt
# Gather PVC status
kubectl get pvc --all-namespaces > /tmp/all-pvcs.txt
# Gather recent events
kubectl get events --all-namespaces --sort-by='.lastTimestamp' > /tmp/events.txt
# Gather cmsdev logs and artifacts
cd /opt/cray/tests/install/logs/cmsdev/<timestamp>
tar -czf /tmp/cmsdev-evidence.tgz cmsdev.log artifacts.tgz
# Gather Python test logs
ls -lt /opt/cray/tests/integration/logs/csm/cmstools/*/If you need to convert legacy logs to new format:
# Using default directory
/usr/local/bin/convert_cmsdev_logs.sh
# Using custom directory
/usr/local/bin/convert_cmsdev_logs.sh /path/to/custom/logdirThis script (see convert_cmsdev_logs.sh):
- Scans for unique run tags
- Creates timestamped directories
- Extracts logs per run tag
- Moves associated artifacts
- Removes original log file
- Infrastructure: Pod failures, PVC issues, node problems
- Service: API errors, functionality failures
- Test: Test logic errors, false positives
- Environment: Configuration, network, permissions
# Gather system state
kubectl get pods --all-namespaces -o wide > /tmp/all-pods.txt
kubectl get pvc --all-namespaces > /tmp/all-pvcs.txt
kubectl get events --all-namespaces --sort-by='.lastTimestamp' > /tmp/events.txt
# Service-specific info
kubectl describe pod <failing-pod> -n services > /tmp/pod-describe.txt
kubectl logs <failing-pod> -n services > /tmp/pod-logs.txt| Category | Common Causes | Resolution Path |
|---|---|---|
| Infrastructure | Node issues, storage problems, network issues | System administration |
| Service | Service bugs, configuration errors, dependency failures | Service team escalation |
| Test | Test assumptions invalid, timing issues, environment changes | Update test code |
| Environment | Wrong versions, missing configuration, RBAC issues | Configuration management |
# Re-run the specific service test
/usr/local/bin/cmsdev test <service> --verbose
# Re-run Python test
/opt/cray/tests/integration/csm/barebones_image_test
/opt/cray/tests/integration/csm/cfs_sessions_rc_test# Check API Gateway
curl -k https://api-gw-service-nmn.local/apis/
# BOS health
curl -k https://api-gw-service-nmn.local/apis/bos/v2/healthz
# CFS health
curl -k https://api-gw-service-nmn.local/apis/cfs/v3/healthz
# IMS version (requires auth)
cray ims versions list
# VCS health
curl -k https://api-gw-service-nmn.local/vcs/api/v1/versionkubectl get pods -n services -o wide # All service pods
kubectl logs <pod-name> -n services # Pod logs
kubectl logs <pod-name> -n services --previous # Previous pod logs (if crashed)
kubectl describe pod <pod-name> -n services # Pod details
kubectl get events -n services --sort-by='.lastTimestamp' # Recent events
kubectl get pvc -n services # PVC status
kubectl exec -it <pod-name> -n services -- /bin/sh # Exec into pod# Check current state
kubectl get deployment -n services cray-cfs-operator
# Restore if scaled to 0
kubectl scale deployment -n services cray-cfs-operator --replicas=1
# If CFS v2 page-size was modified, restore via CLI
cray cfs v3 options update --default-page-size 1000# Find latest cmsdev test run
ls -lt /opt/cray/tests/install/logs/cmsdev/ | head -5
# Search for errors in cmsdev logs
grep -i error /opt/cray/tests/install/logs/cmsdev/<timestamp>/cmsdev.log
# Find latest Python test log
ls -lt /opt/cray/tests/integration/logs/csm/cmstools/barebones_image_test/
ls -lt /opt/cray/tests/integration/logs/csm/cmstools/cfs_sessions_rc_test/