AI agents fail one task in three. We built for that.

The honest number nobody puts on a landing page, and what it changes about how an AI teammate should hand work back to you.

The UnaTask team, Building UnaTask. Aug 2, 2026. 5 min read.

More on AI teammates

Here is a number you will not find on the pricing page of any AI tool, including, until this sentence, ours. The 2026 Stanford AI Index put agents at roughly one failure in every three benchmark tasks. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, mostly on cost and on value nobody could point to.

Sit with that for a second, because it is not the damning statistic it first looks like. Two out of three is a very good hit rate for something that costs a fraction of a cent and takes forty seconds. It is a terrible hit rate for something you are told to trust and forget.

Almost every AI teammate on the market is sold on the second story. Delegate it, walk away, it is handled. Then the third task comes back wrong, quietly, in a place nobody was looking, and the thing that was supposed to save an afternoon costs a week. The failure is not the model being bad. The failure is a product that promised you would not have to look.

Design for the one in three, not the two in three

So we started from the opposite assumption. Assume it will sometimes be wrong, and make being wrong cheap. That single decision shaped everything about how an AI teammate works in UnaTask.

  • The work comes back as a comment on the task, written by the teammate, with any files attached. It does not overwrite anything you wrote.
  • The task then moves to Review. Not Done. Somebody looks before anything counts as finished.
  • Under the work there are three buttons: Good as is, I fixed it, Not usable. Ten seconds, and the record of what happened stays on the task.
  • Nothing starts on its own unless you switch it on. Every project has its own toggle, and a paused teammate stays on the team without picking anything up.
  • There is a ceiling on how many runs a workspace can do in a day, so a loop or a bad template cannot quietly spend your afternoon.

None of that is dramatic. That is the point. When the wrong answer lands in a comment on a task that is sitting in Review, the cost of the one in three is thirty seconds of reading and a click. When the wrong answer lands in a document that says Done, the cost is whatever happens next.

Reviewable beats autonomous

There is a version of this product we could have built where the AI moves your calendar, closes your tasks and files things without asking. It demos beautifully. We have watched people use tools like that, and what actually happens is a slow loss of trust: you start checking everything it did, which is strictly more work than doing it yourself, and then you stop using it.

The teams getting real value out of AI right now are not the ones who handed over the most. They are the ones who put it where a human was already going to look. A draft on the task it belongs to, in front of the person who asked for it, is that place. It is also, not by accident, the only shape where the failure rate stops mattering very much.

And the one in three gets better with use, because a correction is not thrown away. Paste back the version you actually used and that becomes a rule the same teammate carries into its next job. The number is not a fixed property of the model. It is a property of how much your teammate has been told, and by whom.

We do not sell autonomy. We sell work you can look at in thirty seconds and either take or send back.

Read next

  • What you can actually hand to an AI teammate
  • Teach your AI teammate once, and it remembers
  • Assign a task to an AI the way you assign it to a person

More from the blog

  • UnaTask
  • Features
  • Mac app
  • Pricing
  • Blog
  • Comparisons
  • For your team
  • Updates