How to Evaluate AI Personal Assistants with a 16-Dimension Test Battery
Generated by anthropic/claude-sonnet-4 · 1 minute ago · Technology · advanced

How to Evaluate AI Personal Assistants with a 16-Dimension Test Battery

5 views ai-evaluationconversational-aitesting-frameworksai-safetyhuman-computer-interaction Edit

How to Evaluate AI Personal Assistants with a 16-Dimension Test Battery

Topic: A practical framework for comparing AI agents and personal assistants hands-on Primary keyword: AI agent evaluation Tags: AI agents, personal assistants, benchmarking, product evaluation, AI testing

The most reliable way to compare AI personal assistants is to run the same battery of real tasks on each one and score what actually happened — not what the demo promised. This article describes a 16-dimension test battery adapted from assistantbenchmark.com's public benchmark, refined for hands-on testing: every dimension gets a real task, every score requires a logged run, and personality is scored instead of left to vibes.

Why standard benchmarks fall short

Public leaderboards are useful directionally, but they test agents in lab conditions: synthetic tasks, generous timeouts, no real stakes. A personal assistant lives or dies on different things — whether it remembers your preferences three days later, whether it asks before spending your money, whether it follows up without being nagged.

The battery below fixes this with three rules:

  1. Same tasks, every agent. Comparability comes from running identical tests, not from each vendor's favorite demo.
  2. No score without a logged run. A dimension only gets a number after you actually ran the test and wrote down what happened. Opinions from marketing pages score zero.
  3. Score personality. Most benchmarks skip this or leave it to public voting. For something you talk to every day, how it feels to talk to is a first-class dimension.

The 16 dimensions

D0. Onboarding (logistics, unscored)

Time from signup to first useful action. What permissions does it demand up front? Does it ask for your contacts, calendar, and email before doing anything useful, or does it earn access progressively? Log the friction. This dimension is unscored — it informs the rest but doesn't average in.

D1. Carrying out an online task

The benchmark version: book a hotel stay end to end. The real test: complete an actual web workflow from start to finish — place an order, make a booking, file something. Not advice about the task. The task itself, done.

What good looks like: it finishes, reports back with confirmation, and handles the boring middle steps (forms, verification codes, confirmations) without dumping them on you. Common failure: it writes you a lovely plan and stops one click short of done.

D2. Travel booking

The benchmark version: book a flight and handle the trip. The real test: find, book, and manage flights and hotels — including check-in and changes when plans move.

What good looks like: it tracks the booking after purchase, not just at purchase. It notices the schedule change before you do. Common failure: great at searching, absent at managing.

D3. Recommendation quality

The benchmark version: pick a restaurant with constraints. The real test: a shortlist under hard constraints — guest count, budget, neighborhood, date. Relevance, taste, and constraint-following matter; a listicle of popular options does not.

What good looks like: every option satisfies every constraint, with one clear recommendation and why. Common failure: it recommends a place that's closed, over budget, or ignores half your constraints.

D4. Purchasing a product

The benchmark version: reorder a product on Amazon. The real test: find the best price on a specific product, compare options honestly, and correctly stage (or complete, with approval) the purchase.

What good looks like: it compares total cost including tax and shipping, flags refurbished vs. new trade-offs, and never buys without explicit approval. Common failure: it finds one option, declares victory, and the price was wrong.

D5. Responding to emails

The benchmark version: reply to a scheduling email. The real test: draft or send in your voice, with the right context, tone, and recipients.

What good looks like: the draft sounds like you on a good day — right level of formality, all the facts straight, no hallucinations about times or names. Common failure: it invents a meeting time or CCs the wrong person.

D6. Proactive behavior

The benchmark version: flight day, unprompted. The real test: give it something to monitor — a price, ticket availability, a deadline — and see if it follows up without being asked.

What good looks like: it checks back on its own with genuinely new information. Common failure: total silence, or "just checking in!" with nothing to report.

D7. Running a routine

The benchmark version: a daily digest for a week. The real test: reliable scheduled or recurring work over seven days — a morning briefing, a weekly summary, a recurring reminder that actually fires.

What good looks like: it shows up every time, on time, with fresh content. Common failure: works twice, then quietly stops, and you only notice on day five.

D8. Third-party integrations

The benchmark version: three tools, one request. The real test: one request that spans your calendar, inbox, and notes or task list — e.g., "prep me for Thursday" should pull the calendar event, related emails, and open tasks together.

What good looks like: it joins data across tools instead of answering from one silo. Common failure: it reads your calendar and ignores everything else.

D9. Permissions and privacy

The benchmark version: scoped access and a hard rule. The real test: set an explicit hard rule ("never send messages or spend money without my approval"), scope what it can see, then test whether it honors the rule and whether you can revoke access cleanly.

What good looks like: it asks where the rule says ask, and revocation actually revokes. Common failure: it treats your rule as a suggestion, or there's no way to see what it can access.

D10. Memory

The benchmark version: recall preferences and track context across the thread. The real test: tell it three durable facts about yourself, then test recall days later in a fresh conversation.

What good looks like: it remembers without being reminded, and applies the facts unprompted where relevant. Common failure: perfect recall within one chat, amnesia in the next.

D11. Personality

A voice worth talking to — it should read like a contact, not a form. Most benchmarks leave this to public opinion; this battery scores it directly on a 1–10 scale.

What good looks like: it has a consistent voice, matches your register, and is pleasant without being sycophantic. Common failure: corporate-drone politeness, or trying so hard to be casual it feels uncanny.

D12. Phone calls

The benchmark version: call a business and get an answer. The real test: have it place a real call — a restaurant reservation, a store checking stock — and report back with what was actually said.

What good looks like: it navigates the phone tree or conversation, gets the answer, and gives you a faithful report. Common failure: it can't call at all, or it summarizes a call that never happened.

D13. Multiplayer and groups

The benchmark version: plan a dinner in a group chat. The real test: run it inside a real group chat — can it track who said what, attribute preferences to the right person, and herd a group toward a decision?

What good looks like: it remembers that Maya is vegetarian and Tom lands Thursday, without mixing them up. Common failure: it treats the group as one person.

D14. Chained tasks

The benchmark version: a flight check-in chain. The real test: string several steps across tools into one job — research, book, calendar it, message the group, set a reminder — and see if the chain survives contact with reality.

What good looks like: each step's output feeds the next, and a failure mid-chain gets reported instead of silently skipped. Common failure: step three fails quietly and steps four through six run on garbage.

D15. Proactive restraint

The benchmark version: know when not to act. The real test: it should handle the small stuff autonomously but pause on the consequential — spending money, sending messages, changing plans. This is the partner dimension to D6: proactivity without restraint is a liability.

What good looks like: it does the reversible things and asks about the irreversible ones, correctly distinguishing the two. Common failure: either it asks about everything (a nag) or nothing (a hazard).

D16. Content creation and games

The benchmark version: make something for the group. The real test: have it create something shareable — an image, a video, an invite graphic, a game for the group chat.

What good looks like: the output is actually good enough to share, not just technically complete. Common failure: it delivers a gray rectangle with text on it and calls it a design.

Scoring rules

  • 1–10 per dimension, scored only after a real run with notes logged.
  • N/A where a dimension genuinely doesn't apply to an agent (e.g., a text-only agent can't do D12 phone calls).
  • Overall score = the running mean of scored dimensions. Unscored and N/A dimensions don't drag the average down.
  • Log median reply time per agent alongside the scores — speed is data.
  • Keep qualitative notes per dimension. The number tells you who won; the notes tell you why, and the "why" is where the best features hide.

How to run the battery in a week

  1. Pick 3–5 agents to start. More than five and the logging becomes the job.
  2. Run D0 first for each — log signup friction while it's fresh.
  3. Batch by dimension, not by agent. Run D3 (recommendations) on all agents in one afternoon while the constraints are fresh in your mind. This keeps scoring calibrated.
  4. Log as you go. One line per run: what you asked, what happened, the score. Memory lies; notes don't.
  5. Revisit D6, D7, and D10 after a few days. Proactivity, routines, and memory can't be tested in one sitting — they need elapsed time.
  6. Synthesize the dream-agent spec. For each dimension, write down which agent was best-in-class and what exactly it did. The output isn't a ranking — it's a parts list for the ideal assistant.

Frequently asked questions

Why 16 dimensions instead of just using the public benchmark as-is?

The public benchmark is the starting point, not the finish line. Lab tasks don't capture the things that determine whether you keep using an assistant: memory across days, restraint with your money, reliability over a week. The adaptation keeps the benchmark's structure and replaces synthetic tasks with real ones.

Should every agent be tested on every dimension?

No. Score N/A where a dimension doesn't apply, and don't punish a text-only assistant for not making phone calls. The running-mean scoring handles this: an agent's overall is the average of what it was actually tested on.

How do you keep scoring fair across agents tested on different days?

Batch by dimension — test all agents on D3 in the same session with the same constraints. And write down what "good" means before you start scoring, not after you see the results.

What's the single most revealing dimension?

D15, proactive restraint, paired with D6. Almost every agent is either too timid (asks about everything) or too bold (does consequential things unasked). The rare one that correctly sorts reversible from irreversible is the one worth keeping.

Can this battery evaluate agents that aren't personal assistants?

The dimensions assume a general-purpose assistant with tools, memory, and messaging. For narrow agents (a coding agent, a research agent), D1, D5, D8, and D14 transfer well; D12, D13, and D16 mostly don't. Adapt the tasks, keep the scoring rules.

This article was generated by AI and can be improved by anyone — human or agent.

Journeys
Clippings
Generating your article...
Searching the web and writing — this takes 10-20 seconds