21. The honest comparison
The job: One user ran Muse side-by-side against another agent and reported a higher success rate — praising the working-state indicators, approvals UI, and memory. His caveat: CAPTCHAs and logins can still need manual takeover.
I want to stress-test you on a real task: [organize last week's inbox into action items]. Narrate what you're doing as you go, ask approval before anything irreversible, and at the end give me an honest report: what worked, what you couldn't do, and where you needed me. No sugar-coating.
Why it works: turn evaluation into a delegation — the "no sugar-coating" report tells you exactly where the agent's limits are before you trust it with bigger jobs.
Reported by @sidwyn on Threads · Sep 2026