How DevClocked Measures
How DevClocked classifies activity with confidence scores, computes leverage (measured vs estimated), scores flow and focus, captures token cost exactly, and where the limits are.
Every number DevClocked shows is derived, not guessed. Here is exactly how each metric is built, when it is measured versus estimated, and where the limits are.
Session, dev time, and leverage
Three words do most of the work in this product, and they are easy to mix up. This is what each one means.
- Session is the container: the wall-clock time you were present in a dev surface, whether that is an editor, a CLI plugin, Claude Code or another agent CLI, a tracked site in the browser extension, or Figma. It is capped at the wall clock for each of your local days, so two machines running at once never bill the same minute twice.
- Dev time means the minutes inside a session where code was actually being written, by you or by an agent, counted as a union rather than added up. It can never be larger than the session it sits in. Nothing measures it yet, so the web app no longer shows a Dev time label; the desktop app still does until its next release.
- Leverage is a multiplier: your attended hours plus the agent hours running inside them, including parallel agents, divided by your attention. Attention is the session hours, discounted by up to 55% when the session was mostly orchestrating agents. Above 1x the number is compressed by a log curve and scaled by a quality credit, so a low-quality session pulls the score down, a very high-quality one lifts it slightly, and a fan-out of subagents cannot run the number away with you.
Worked example. A 1 hour session in which agents ran for 3 hours would be 4x if nothing were dampened. The shipped score lands between roughly 3.4x and 5.7x, depending on session quality and how orchestration-heavy the session was. The Trend chart's Hands-on line is a different cut again: session minutes with no agent running. It shows you the shape of the day, and it is not an input to the leverage score.
Time Slice classification
Activity ticks are grouped into blocks and each block is labelled as one of six categories. Each classification carries a confidence from 0.3 to 0.95 based on how directly the signals matched — a live debug session scores high, a default file-edit guess scores low. That confidence is shown on every block and averaged per session.
- Building — file creation, significant additions, source edits
- Debugging — debug sessions, breakpoints, terminal errors, test files
- Refactoring — refactor commands, balanced add/delete churn
- Planning — docs, AI tools, design tools, issue trackers
- Config — config files, CI paths, dependency manifests, dotfiles
- Review — read-heavy passes over code, PR/MR review URLs
Leverage
Leverage expresses how much output your attention unlocked. It is computed from the composition of your work — hands-on coding vs agent orchestration — together with agent output signals. When agent runtime telemetry corroborates it, we treat it as measured; without that telemetry it is an estimate from work patterns, and the session view flags it est.
- Measured — agent runtime was captured for the session. Leverage reflects real agent-to-attention ratio.
- Estimated — no agent runtime present. Leverage is inferred from orchestration share and quality — directional, not exact.
Leverage is decomposable: it always drills into its causes rather than standing alone as a bare multiplier.
Flow & focus
Flow and focus are engagement signals, not verdicts.
- Flow (0–100) — rewards consistent tick intervals and focused file spread; penalises context switching. High flow means sustained, uninterrupted work.
- Focus (0–100) — rewards write activity and time spent in your top files. High focus means engaged, file-sticky work rather than scattered attention.
Token cost
Usage is read from the model's own top-level token counts — never summed across intermediate iterations, which would double-count. Each usage row is priced against a version-pinned price book, so costs are exact for priced models. Usage from models not yet in the book is flagged as unpriced rather than silently valued at zero.
Agent efficiency
For sessions with agent turns we compute two signals: tokens per turn, and a cache-hit ratio (cached tokens over input-side tokens) that proxies context resets — a low ratio means each turn re-sends context instead of reusing it. A run is flagged churny when it produces few lines per turn and its cache-hit ratio is low, and efficient when output per turn is high. When a signal's ingredient is missing it stays blank rather than reading as zero.
Known limits
- Idle handling — short gaps split activity blocks, and sustained silence ends the session — so very deliberate think-time can read as idle.
- Inferred sessions — sessions reconstructed from commits (no live tracker) lack tick-level signal — timing and classification are approximate.
- Estimated leverage — without agent runtime, leverage is directional, not a measured ratio.
- Low-confidence blocks — some blocks classify on weak signals; the confidence score tells you which.
A note on people
These numbers are not for ranking people. They explain how work happened and help you plan capacity — not to stack-rank, review, or police the humans doing it. Read our measurement ethics for the full stance, and what DevClocked collects for exactly which data leaves your machine.
Was this page helpful?