Most teams can tell you whether a new tool feels useful within a week. Whether it's actually worth what they're paying for it is a different question, and it's the one worth answering with numbers instead of a gut check — especially before deciding whether to expand from one division to Full Suite.
The four numbers that matter
- Approval rate over time — what share of an agent's drafted actions get approved as-is, edited, or rejected, and whether that ratio is improving week over week.
- Time-to-first-value — how long between connecting a division and the first agent output a human actually used, not just reviewed.
- Hours reclaimed — the honest version of this is self-reported by the team doing the work, cross-checked against what's actually shipping through the approval gate.
- Audit trail incidents — how many logged actions needed a correction after the fact. Zero forever is a red flag that approvals are being rubber-stamped, not a sign of perfection.
What good looks like in month one
As described in the first 30 days with STIV, month one is mostly the connect and learn stages — approval rates should be climbing but still well below where they'll settle, and most of a team's interaction with the system should be reviewing and correcting drafts rather than rubber-stamping them. If approval rate is already near 100% in week two, that's usually a sign the team isn't reading closely yet, not that the agent is unusually good.
What a division that isn't working looks like
The clearest signal isn't a low approval rate — a low rate that's climbing is normal early on. It's a flat or declining approval rate past week four, combined with a team that's stopped engaging with drafts closely because they've learned not to trust them. That combination means the agent isn't learning from the correction signal it's being given, and it's worth a direct conversation about playbooks and data access before assuming the model itself is the problem.
None of this requires a dashboard you don't already have. The audit trail every action already generates — logged, timestamped, and reversible — is the same data you need to answer whether a division has earned its keep. The discipline is in actually looking at it monthly, not waiting for renewal to ask the question.