Research

What agents can actually do, measured rather than asserted.

Evaluation work, benchmarks and notes from Mapier Labs. Results appear here when they are real; a slug is published before the numbers so the first link stays the right one.

  1. BenchmarkIn progressA benchmark for agents that hold a conversation over daysMost agent evaluations end when the task does. Real messaging does not: the agent is interrupted, contradicted, and asked to remember something from last week. We are building the evaluation for that.