← All 28 cases

💼 Work

21. The honest comparison

The job: One user ran Muse side-by-side against another agent and reported a higher success rate — praising the working-state indicators, approvals UI, and memory. His caveat: CAPTCHAs and logins can still need manual takeover.

I want to stress-test you on a real task: [organize last week's inbox into action items]. Narrate what you're doing as you go, ask approval before anything irreversible, and at the end give me an honest report: what worked, what you couldn't do, and where you needed me. No sugar-coating.

Why it works: turn evaluation into a delegation — the "no sugar-coating" report tells you exactly where the agent's limits are before you trust it with bigger jobs.

Reported by @sidwyn on Threads · Sep 2026

Was this useful?

Want templates instead of full scenarios? Open the prompt library →