Building AgentMeasure — open measurement & conformance infrastructure for AI agents.
Measure what agent telemetry actually counts: attempts, operations, retries, token accounting, and evidence boundaries.
AgentMeasure · Conformance vectors · Website · 中文主页
- ✓ OpenLIT — token-accounting invariant merged upstream (reasoning ⊂ output; adding them reports tokens that never happened)
- ✓ Urusilla — 2 external conformance vectors → 3 defects found and fixed, each with the finder's fixture as the regression guard
- ✓ Agent semantics reviewed in public across Langfuse · DeepEval · OTel GenAI
AgentMeasure tests whether agent telemetry metrics actually mean what their labels claim.
- Attempts vs operations — one logical use vs N executions
- Retry accounting — a retried call is not two users
- Token double counting — subsets must not be added into totals
- Evidence vs inference — what a trace proves vs what it merely suggests
- Eval repeatability — n runs are n measurements, not one verdict
Proven in the wild: an invariant merged into OpenLIT's main, two external conformance vectors from an independent contributor, three externally found defects fixed publicly. Long term: a common measurement layer for the Agent Capability Economy.
Try the 2-minute local demo · Send a trace, get a measurement check
Product and industry research in four directions. Facts first; conclusions revised as evidence changes.
| Direction | Central question |
|---|---|
| Embodied Intelligence | how robots move from demo to real tasks, product definition, and commercialization |
| Multimodal Interaction | what the next interaction paradigm looks like when voice, vision, and space become I/O |
| On-device AI | how models, compute, privacy, and form factors jointly shape the AI product experience |
| Agents | how agents acquire context, choose capabilities, call software — and form a new software economy |
Judgments worth keeping, discussing, and falsifying (中文):
- AI Native 之后,产品的基本单位变了
- RAG 之后,Agent 需要 Context Recommendation
- 空间计算需要自己的“触控时刻”
- 家庭机器人为什么技术上更难,经济上可能更丰富
| Tool | What it does |
|---|---|
| User Demand Research (SURE) | auditable demand evidence and research reports from goals, scope, and sample design (Skill + CLI + MCP) |
| iRead Research Monitor | continuous source discovery with evidence-aware daily / weekly / monthly digests |
| Bilibili Video to Transcript | public videos → timestamped, searchable research text |
| Roy's AI Product Research Library | public research, essays, and an agent-readable index |
Machine-readable tool catalog (agent routing): https://raw.githubusercontent.com/roy-tong/roy-tong/main/agent-tools.json
Router skill:
gh skill preview roy-tong/roy-tong find-research-tool
gh skill install roy-tong/roy-tong find-research-tool --agent codex --scope userProduct manager and repeat founder, ~10 years across AI, software, and intelligent hardware. About · GitHub · X / Twitter · Contact

