Investigating Benchmark Performance

SkillMonitoring & ops

Investigating benchmark performance by running Golem benchmarks with OTLP tracing enabled and analyzing trace data from Jaeger. Use when diagnosing slow benchmarks, understanding benchmark call flows, or profiling span durations during benchmark runs.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Investigating Benchmark Performance skill

What this skill tells your AI

The instructions your AI receives, as published by golemcloud/golem in .agents/skills/investigating-benchmark-performance/SKILL.md and read by ahel’s review.

Run Golem benchmarks with distributed tracing enabled and analyze the resulting traces in Jaeger to understand performance characteristics, call flows, and bottlenecks.

Prerequisites

  • Docker and Docker Compose installed
  • The monitoring stack defined in integration-tests/monitoring/docker-compose.yml
  • The benchmark WASM components built with cargo make build-benchmark-components (see Step 1.5)
  • The Golem service binaries built with release profile: cargo build --release -p golem-worker-executor ... (see Step 1.5)
  • The benchmark runner binary built with benchmarks profile: cargo build --profile benchmarks -p integration-tests --bin benchmarks (see Step 1.5)

Step 1: Start the Monitoring Stack

From the integration-tests/monitoring/ directory:

cd integration-tests/monitoring
docker compose down && docker compose up -d

This starts:

  • Jaeger — UI on localhost:16686, OTLP collector on localhost:4318
  • Prometheus — on localhost:9090
  • Grafana — on localhost:3000 (admin/admin)

Step 1.5: Build Binaries

Build the benchmark components and both sets of native binaries from the repository root.

Benchmark components

cargo make build-benchmark-components

This builds the WASM components selected by the benchmarks component group. Rebuild them when their sources, SDKs, WIT inputs, or component build tooling change.

Service binaries (release profile)

The benchmark runner in spawned mode starts Golem services as child processes. It expects pre-compiled release binaries at target/release/.

cargo build --release \
  -p golem-component-compilation-service \
  -p golem-worker-service \
  -p golem-worker-executor \
  -p golem-shard-manager \
  -p golem-registry-service

You must rebuild after every code change — the benchmark runner uses whatever binaries are on disk. To rebuild only the crate you changed:

cargo build --release -p <crate-you-changed>

Benchmark runner binary (benchmarks profile)

The benchmark runner itself must be built with the benchmarks profile (defined in root Cargo.toml, inherits from release with panic = "unwind"):

cargo build --profile benchmarks -p integration-tests --bin benchmarks

This produces the binary at target/benchmarks/benchmarks.

Step 2: Run Benchmarks with OTLP Tracing

The benchmarks binary is at integration-tests/src/benchmarks/all.rs and is built as the benchmarks binary from the integration-tests crate. It has a built-in --otlp flag that configures all spawned Golem services to export traces.

--otlp exports at info. To change what is exported, set GOLEM_OTLP_FILTER — it takes RUST_LOG syntax and is used exactly as written, so GOLEM_OTLP_FILTER=info,h2=off quietens one crate and debug widens everything. The filter in force is logged at startup as otlp_filter. See the investigating-executor-performance skill for the target names.

This traces internal Golem service operations. It does not enable or inspect the Golem OTLP exporter plugin, which exports agent/runtime telemetry through the separate docker-examples/otlp-collector stack.

Running a single benchmark

The CLI takes the benchmark name as a positional argument, followed by the spawned subcommand (which tells it to spawn services locally). The --build-target defaults to target/release which is the correct path for the service binaries.

Run the benchmark using the benchmarks-profile binary:

./target/benchmarks/benchmarks \
  --otlp \
  benchmark \
  --iterations <N> \
  --cluster-size <S> \
  --size <W> \
  --length <L> \
  <benchmark-name> \
  spawned

Note: --size and --cluster-size accept multiple values by repeating the flag (e.g., --size 1 --size 10), not comma-separated.

Available benchmarks

NameDescriptionRecommended quick-run parameters
cold-start-unknown-smallFirst-time invocation of a never-instantiated small component--size 1 --size 5 --size 10 --length 2 --cluster-size 1 (with and without --disable-compilation-cache)
cold-start-unknown-mediumFirst-time invocation of a never-instantiated medium component--size 1 --size 5 --size 10 --length 5 --cluster-size 1 (with and without --disable-compilation-cache)
latency-smallCold and hot invocation latency for a small component--size 100 --size 500 --size 1000 --length 2 --cluster-size 1
latency-mediumCold and hot invocation latency for a medium component--size 100 --size 500 --size 1000 --length 5 --cluster-size 1
sleepMeasures sleep/suspend overhead--size 10 --size 100 --size 500 --length 10000 --cluster-size 1
durability-overheadMeasures the overhead of durable execution--size 10 --size 100 --size 1000 --length 5000 --cluster-size 1
throughput-echoThroughput benchmark with echo workload--size 1 --size 10 --size 100 --length 1000 --cluster-size 1 --cluster-size 5
throughput-large-inputThroughput benchmark with large input payloads--size 1 --size 10 --length 100 --length 10000 --length 100000 --cluster-size 1 --cluster-size 5
throughput-cpu-intensiveThroughput benchmark with CPU-intensive workload--size 1 --size 10 --length 100 --cluster-size 1 --cluster-size 5

Example: single benchmark with tracing

./target/benchmarks/benchmarks \
  --otlp \
  benchmark \
  --iterations 1 \
  --cluster-size 1 \
  --size 100 \
  --length 2 \
  latency-small \
  spawned

Running a benchmark suite

Benchmark suites are YAML files in integration-tests/benchmark_suites/. They define multiple benchmarks with their parameters.

./target/benchmarks/benchmarks \
  --otlp \
  suite \
  integration-tests/benchmark_suites/quick-all.yaml \
  spawned

Available suites:

  • quick-all.yaml — All benchmarks with reduced parameters, suitable for quick local runs
  • ci.yaml — CI configuration

Suite results can be saved to JSON:

./target/benchmarks/benchmarks \
  --otlp \
  suite \
  --save-to-json tmp/benchmark-results.json \
  integration-tests/benchmark_suites/quick-all.yaml \
  spawned

Additional CLI flags

FlagDescription
--otlpEnable OTLP tracing for all spawned services
--jsonOutput results as JSON instead of human-readable format
--primary-onlyOnly display primary results (no per-worker breakdown)
--quietReduce log verbosity to warnings only
--verboseIncrease log verbosity to debug level

Benchmark parameters

ParameterMeaning
iterationsNumber of times to repeat the benchmark
cluster-sizeNumber of worker executors in the test cluster
sizeWorkload size (e.g., number of workers)
lengthWorkload length (benchmark-specific, e.g., payload size, iteration count)
disable-compilation-cacheDisable compilation caching to measure cold compilation

Step 3: View Traces in Jaeger

Open http://localhost:16686 in a browser. The benchmarks service name depends on the spawned services — look for service names like worker-executor, worker-service, component-compilation-service, etc.

Step 4: Analyze Traces via Jaeger API

Jaeger exposes an HTTP API at localhost:16686. Use it to programmatically analyze trace data.

Fetch traces

# List services (to find the correct service names)
curl -s 'http://localhost:16686/api/services' | python3 -m json.tool

# Fetch traces for a specific service
curl -s 'http://localhost:16686/api/traces?service=worker-executor&limit=1000&lookback=1h' \
  -o tmp/traces.json

Analysis patterns

All examples assume traces are saved in tmp/traces.json. Span durations in the Jaeger JSON are in microseconds (divide by 1000 for milliseconds).

Summary statistics
python3 -c "
import json
data = json.load(open('tmp/traces.json'))
traces = data['data']
total_spans = sum(len(t['spans']) for t in traces)
print(f'Traces: {len(traces)}, Total spans: {total_spans}')
"
Operation name distribution
python3 -c "
import json
from collections import Counter
data = json.load(open('tmp/traces.json'))
ops = Counter()
for t in data['data']:
    for s in t['spans']:
        ops[s['operationName']] += 1
for op, count in ops.most_common(30):
    print(f'{count:6d}  {op}')
"
Find slowest spans
python3 -c "
import json
data = json.load(open('tmp/traces.json'))
spans = []
for t in data['data']:
    for s in t['spans']:
        spans.append((s['duration']/1000, s['operationName'], s['traceID'][:12]))
spans.sort(reverse=True)
for dur_ms, op, tid in spans[:20]:
    print(f'{dur_ms:10.1f}ms  {op}  trace:{tid}')
"
Find error spans
python3 -c "
import json
data = json.load(open('tmp/traces.json'))
for t in data['data']:
    for s in t['spans']:
        for tag in s.get('tags', []):
            if tag['key'] == 'otel.status_code' and tag['value'] == 'ERROR':
                dur = s['duration'] / 1000
                print(f'{dur:.1f}ms  {s[\"operationName\"]}  trace:{s[\"traceID\"][:12]}')
"
Trace size distribution (spans per trace)
python3 -c "
import json
from collections import Counter
data = json.load(open('tmp/traces.json'))
sizes = Counter(len(t['spans']) for t in data['data'])
for size, count in sorted(sizes.items()):
    print(f'{count:4d} traces with {size:4d} spans')
"
Inspect single-span traces

Handed-off work starts a new root trace linked to its producer, so a single-span trace is normally expected rather than an orphan. In particular, invocation execution is intentionally separate from the request trace: it is linked to the producing enqueue_invocation span, and both sides carry the same idempotency_key. Use the OpenTelemetry link and idempotency_key to correlate them; do not expect a parent-child relationship across the hand-off.

Investigate a single-span trace only when it has neither an expected link nor a connected parent. That is a context-propagation gap.

python3 -c "
import json
from collections import Counter
data = json.load(open('tmp/traces.json'))
single_spans = Counter()
for t in data['data']:
    if len(t['spans']) == 1:
        single_spans[t['spans'][0]['operationName']] += 1
print(f'Total single-span traces: {sum(single_spans.values())}')
for op, count in single_spans.most_common(15):
    print(f'{count:4d}  {op}')
"
Identify background noise traces

Background operations are bounded to one tick or operation, but their frequency can dominate trace volume. Filter them out for focused analysis:

python3 -c "
import json
data = json.load(open('tmp/traces.json'))
NOISE = {'oplog_background_transfer', 'ephemeral_oplog_background_transfer',
         'oplog_forwarding_flush', 'oplog_forwarding_threshold_flush',
         'scheduler_tick', 'quota_renewal',
         'resource_limits_batch_update', 'agent_status_flush_sweep'}
clean = [t for t in data['data']
         if not any(s['operationName'] in NOISE for s in t['spans'])]
print(f'Total: {len(data[\"data\"])}, After filtering noise: {len(clean)}')
"

Known Caveats

  • Multiple services: Unlike worker-executor tests (which run in-process), benchmarks spawn separate Golem services. The --otlp flag configures all spawned services to export traces. Look for multiple service names in Jaeger.
  • Trace context propagation: gRPC calls within a request path propagate context via traceparent headers. Handed-off work is intentionally a separate linked root; correlate an invocation's producer and execution with its OpenTelemetry link and shared idempotency_key. If a trace has neither a normal parent nor a link, verify the monitoring stack was running before the benchmark.
  • Span queue size (OTEL_BSP_MAX_QUEUE_SIZE): The BatchSpanProcessor has a default queue size of 2048 spans. Under high-throughput benchmarks this queue can overflow, causing spans to be silently dropped (logged as BatchSpanProcessor dropped a Span due to queue full). The test framework sets OTEL_BSP_MAX_QUEUE_SIZE=262144 for spawned services when --otlp is enabled. Set the same value for the benchmark runner if it emits spans: OTEL_BSP_MAX_QUEUE_SIZE=262144 ./target/benchmarks/benchmarks --otlp ....
  • Background loop noise: Background spans close after each tick or operation; they are not benchmark-long spans. Their volume can still obscure benchmark traces.
  • Fresh Jaeger: Always restart Jaeger with docker compose down && docker compose up -d before a new investigation to avoid mixing traces from different runs.
  • Build time: Service binaries must be built with --release. The benchmark runner binary must be built with --profile benchmarks. After any code change, rebuild the affected service binary with cargo build --release -p <crate> before re-running the benchmark.

Resetting Between Runs

cd integration-tests/monitoring
docker compose down && docker compose up -d

This clears all stored trace data so the next benchmark run starts fresh.

Signals

GitHub stars
2k
Forks
212
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
investigating-benchmark-performance
Source
github.com/golemcloud/golem