How measurement works
- Deterministic sampling. Whether a call is sampled is a pure function of its
traceId— anyone can recompute the decision. No randomness, ever. - Cache dimension. Executions served from the compiled-rule cache (
HIT) and executions that had to load + compile first (MISS— first call, right after a deploy) are kept as separate distributions by design. Mixing them would hide both signals. - 60-second windows. Observations aggregate into 60s windows. Missing windows are shown as genuine gaps — never interpolated.
- Honest percentiles. Every percentile is accompanied by its sample count
n. A percentile is shown only when at least 3 observations sit beyond it (n×(1−q) ≥ 3— p50 from 6, p95 from 60, p99 from 300 samples); otherwise it is withheld (nullin the API,–in the console) instead of fabricated.
Slow-rule judgment — relative only
A rule is flagged when its p50 is at least 10× the median of per-rule p50s within the same group, version, phase, and cache state. Judgment requires a cohort of at least 3 rules withn ≥ 100 each. This judgment gate is separate from the per-percentile display gate — a rule
can be flagged while its p99 is still withheld.
Two deliberate choices:
- No absolute thresholds. “Slow” depends on your domain — 50ms may be fine for a batch pipeline and unacceptable for checkout. LexQ never hard-codes a millisecond limit, anywhere.
- Median, not mean. With rules at 1, 1, 1, 1 and 40ms, the mean (8.8ms × 10 = 88ms) lets the 40ms rule hide behind the average it inflated. The median (1ms × 10 = 10ms) catches it.
Console
Open a policy group and select the Performance tab.- Presets 1h / 24h / 7d, a HIT / MISS cache toggle, and per-percentile colors (p50 / p95 / p99) used consistently across cards, tables, and charts.
- The rule table shows samples, percentiles, the multiple vs. baseline (
×1.2), and a Slow badge when flagged. - Click a rule for the detail view: a summary per phase × cache state (merged over the window), plus a dot-per-window time series. Dots are raw window values; empty stretches mean nothing ran.
API
Two read endpoints under the partner API (x-api-key auth):
versionId (default: live version), from / to
(ISO-8601 instants, half-open range [from, to), default: last 24h), and cacheState
(HIT | MISS, default HIT, overview only). An empty or inverted range (from ≥ to)
returns error P-035.
Response semantics worth knowing:
- A percentile field is
nullwhen its display gate (n×(1−q) ≥ 3) is not met;nis always present. Mixed responses — p50 present, p99null— are normal. baselines[].statusisINSUFFICIENT_COHORTuntil 3 qualifying rules exist.droppedRowscounts corrupted histogram rows that were skipped — losses are counted, never hidden.
CLI & MCP
lexq_profile_overview and
lexq_profile_rule.
Overhead & retention
Profiling is designed to never compete with your decisions:- Measured overhead at the default 1% sampling: +0.12% on engine p50 (JMH, 20-rule fixture). Instrumentation failures are isolated and counted — they never propagate to the decision path.
- Windows are queryable for 35 days, then archived. Available on every plan.
Per-rule detail comes from a 1% sample, so low-traffic groups accumulate rule-level
percentiles slowly — the group total is recorded on every call regardless. A
– is the
system refusing to guess, not a missing feature.
