From ec3c30b5b2b88fc80d64a8e09972f587c454fa29 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Jordan=20Aug=C3=A9?= Date: Fri, 31 Jul 2026 14:18:55 +0200 Subject: [PATCH] chore(report): arwg - july 2026 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Signed-off-by: Jordan Augé --- reporting/2026-07-report.md | 83 +++++++++++++++++++++++++++++++++++++ 1 file changed, 83 insertions(+) create mode 100644 reporting/2026-07-report.md diff --git a/reporting/2026-07-report.md b/reporting/2026-07-report.md new file mode 100644 index 0000000..df96ac2 --- /dev/null +++ b/reporting/2026-07-report.md @@ -0,0 +1,83 @@ +# [Report] Accuracy & Reliability – July 2026 + +**Working Group:** WG Accuracy & Reliability + +**Reporting period:** July 2026 + +**Chair(s):** Jordan Augé (Cisco) & Casper Nielsen (Diagrid) + +**Date submitted:** 2026-07-31 + +--- + +## Summary + +- D1 (terminology/methodology support deliverable) moved from initial Google Docs drafting to a first complete GitHub PR draft for open review and iterative updates. +- Based on D1, ARWG contributed an initial set of terms to the shared Taxonomy & Landscape sheet. +- ARWG narrowed terminology work into staged review batches. +- D2 (enterprise gap-analysis deliverable) evolved from benchmark inventory toward practitioner-priority analysis (personas, requirements, and practical gaps that block direct reuse by practitioners). + +## Progress Against Objectives + +- **Taxonomy / cross-WG alignment** + - ARWG representatives participated in cross-WG meetings. + - ARWG contributed an initial set of terms (based on D1) to the shared Taxonomy & Landscape sheet. + - Submitted terms include ARWG-specific terminology, terms better handled by other WGs, and generic/common terms in the broader taxonomy; feedback in meetings asked for a tighter top-10 priority set because term-by-term discussion can be long. +- **Terminology execution model** + - Internal work was split into 3 milestones with an initial priority subset. Focus was on terms owned by the WG. + - Scope-control evidence: M1 currently has 41 terms that are central to the D1 structure; 23 are the main focus terms (WG-specific, not well defined or conflictual in literature, or ambiguous), while the rest are more trivial/common terms. + - The focus terms cover different accuracy/reliability dimensions used to structure D1, such as outcome correctness, consistency across runs, robustness to perturbations, and operational reliability in real systems. + - Review rule: central parent sections are handled first; deeper leaves are addressed in later milestones. + - Milestones and progress tracking (Google Docs): +- **D1** + - D1 was used as supporting methodology deliverable (approach + literature grounding for terminology). + - Initial draft was developed in Google Docs; first complete draft was submitted as a GitHub PR. + - Objective: continue in-repo iteration and elicit broader structured feedback. +- **D2** + - Scope reframed from benchmark survey to enterprise-focused gap analysis. + - Early finding: disconnect between available research/assets and directly reusable practitioner material. + - Current prioritization framing was based on personas, use cases, and pain points; concrete comparative data remained limited, which motivated the survey. + - Working cadence: an on-demand D2 sync (including survey work) ran in alternation with the main WG meeting. +- **I1 (liaison note)** + - Internal liaison document drafted. + - First contact established with Observability WG. + - Volunteers are being solicited for additional top-priority WG liaisons. + +## Blockers and Risks + +- **Overall feedback volume remains low across work items (D1, D2, taxonomy, and related reviews)** + - Risk: slower convergence on canonical terms, weaker cross-WG adoption confidence, and slower prioritization closure. + - Seasonal effect: summer period and PTOs reduced review bandwidth and writing capacity. + - Note: despite low feedback volume overall, D1 already proved useful as input for D2 framing; D2 discussions gathered a small contributor group, but there are still few active writers. +- **D2 use-case grounding is not mature enough yet** + - Risk: priorities remain weakly anchored without a convincing shared practitioner use-case set. + - Constraint: contributors are at very different adoption stages (early vs very focused exploration), with different focus priorities (for example: cost, quality, and reliability). + - Mitigation in progress: prepare a short survey to rank pain points and converge on a top-priority use-case set. + +## Decisions Needed from the TC + +- Confirm whether the TC recommends engagement mechanisms beyond PR-based review (PR flow improves traceability but may not significantly increase participation on its own). +- Validate D2 direction: practitioner-grounded prioritization (personas + enterprise requirements + actionable gaps), not benchmark-only analysis. +- Approve a short D2 survey to rank pain points and lock top use cases for the next cycle. + +## Next Month's Focus + +1. Following D1 GitHub PR, actively solicit review comments (2w deadline for first iteration). +2. Finalize first milestone term subset and continue global-term promotion proposals, including a reduced top-10 discussion set. +3. Run targeted survey to prioritize top practitioner pain points and use cases for D2 (finalize survey beg Sept) +4. Expand liaison coverage beyond Observability WG for top-priority interfaces. + +## Metrics / Links + +- D1 reference: [D1-terminology-taxonomy/README.md](../D1-terminology-taxonomy/README.md) +- D2 reference: [D3-gap-analysis/README.md](../D3-gap-analysis/README.md) +- Liaison note: [taxonomy-landscape/taxonomy-wg-internal-liaison.md](../taxonomy-landscape/taxonomy-wg-internal-liaison.md) +- D1 GitHub PR link: to be added when published (planned in early August). +- D1 draft (Google Docs): +- D2 gap analysis (Google Docs): +- D2 survey draft (WIP): +- Liaison bootstrap (Google Docs): +- Taxonomy internal milestones/progress (Google Docs): +- Main running notes: +- Cross-WG Taxonomy sheet: +- Meeting notes and top-10 reduction feedback are captured in the main running notes (link above).