Skip to content

Reduce instrumentation test shard duration - #12754

Draft
cgcote wants to merge 1 commit into
masterfrom
ci/rebalance-instrumentation-shards
Draft

cgcote wants to merge 1 commit into
masterfrom
ci/rebalance-instrumentation-shards

Conversation

@cgcote

@cgcote cgcote commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • increase the blocking AMD64 instrumentation test matrix from 8 to 12 shards
  • assign instrumentation projects with a deterministic longest-processing-time allocator
  • seed the allocator with CI Visibility durations for the known heavy projects and a conservative fallback for the rest
  • scope duration weighting to this 12-way job; ARM64 and other matrices retain normal hashing

Motivation

CI Visibility showed that 12-way hash sharding reduced the slowest instrumentation shard from 37m49s to only 35m07s, with a 2.9x spread between the slowest and fastest shards. Moving JDBC alone shortened its old shard but merely moved the critical path to another shard.

This experiment separates the four dominant workloads (Armeria, JDBC, Spring WebFlux, and Lettuce) and balances the remaining measured projects by expected duration.

Validation

  • ./gradlew :buildSrc:spotlessCheck
  • all 12 instrumentation shards select tasks successfully in Gradle dry runs
  • Armeria, JDBC, Spring WebFlux, and Lettuce resolve to distinct shards
  • git diff --check

Notes

This PR contains one CI-only commit based on master.

@cgcote cgcote added comp: testing Testing tag: no release notes Changes to exclude from release notes type: refactoring tag: ai generated Largely based on code generated by an AI or LLM labels Oct 6, 2026
@datadog-datadog-prod-us1-2

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 14.86 s 14.79 s [-0.4%; +1.3%] (no difference)
startup:insecure-bank:tracing:Agent 13.75 s 13.91 s [-1.9%; -0.4%] (maybe better)
startup:petclinic:appsec:Agent 17.25 s 17.16 s [-0.4%; +1.5%] (no difference)
startup:petclinic:iast:Agent 16.93 s 17.05 s [-1.5%; +0.1%] (no difference)
startup:petclinic:profiling:Agent 16.60 s 16.75 s [-2.1%; +0.4%] (no difference)
startup:petclinic:sca:Agent 17.25 s 16.99 s [+0.5%; +2.4%] (maybe worse)
startup:petclinic:tracing:Agent 16.11 s 16.19 s [-1.4%; +0.4%] (no difference)

Commit: cdc2b15c · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@cgcote
cgcote force-pushed the ci/rebalance-instrumentation-shards branch from 230d7c4 to 42970b7 Compare October 6, 2026 11:30
@cgcote
cgcote force-pushed the ci/rebalance-instrumentation-shards branch from 42970b7 to cdc2b15 Compare October 6, 2026 13:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: testing Testing tag: ai generated Largely based on code generated by an AI or LLM tag: no release notes Changes to exclude from release notes type: refactoring

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant