Imagine a folder with 200 PDF invoices that you need in a spreadsheet. An agent with authorised access can read them and prepare a file for review. Some chats can do this too: what matters is the available tools and which steps they can carry out and check.
Choose a specific task, clear limits and a way to check the result before using it.
AI agent vs chatbot: the difference
| What changes | Chat | Agent |
|---|---|---|
| Who does the work | You and the chat tools | The agent, and you review |
| What it can touch | Shared content and authorised connections | Files and services with permission |
| How long it runs | Depends on the task and tools | Depends on the task and limits |
| How it ends | With an answer or a file | With a result to review or unfinished work |
How an AI agent works
An agent combines a model with tools and a work loop. It observes the state, chooses a step, acts and checks the outcome. It can correct mistakes it detects, but it can also miss errors, get stuck or hit a limit. Saying “finished” does not prove the work is correct.
With the invoices, it would look roughly like this:
- Look: it opens the folder. There are 203 files and 3 aren’t invoices.
- Think: better to read them in batches and note what fails.
- Act: it pulls out the date, supplier, net amount, VAT and total of each one.
- Check: it compares amounts with the PDFs, corrects two readings and flags two that remain uncertain. It includes each row’s source file, counts and sums for you to review.
The fourth step helps detect mistakes, but it can fail too. Review the checks and the source documents.
What an AI agent can do today
METR mainly measures computing tasks: it estimates the human expert time for tasks an agent completes at a given success probability. A 50% success rate is neither a guarantee nor a duration of continuous operation. The historical trend in its 2025 study does not prove you can delegate a morning of office work. Test your own task, measure errors and include review time.
AI agent examples for your work
Look for something you do often, that is boring and that is easy to check. For example:
- Tidy up a folder. It proposes how to split the files into subfolders and waits for your go-ahead before moving anything.
- Put invoices into a spreadsheet. Date, supplier and amounts, checked against the source document with missing data identified.
- A one-page report. With the data from that sheet: the month’s spending, the top five suppliers and a chart.
The second one, ready to copy:
Read the fictional PDF invoices in this folder and create invoices.xlsx with the source file, date, supplier, tax ID, net amount, VAT amount and document total. Check each amount against the PDF. Record multiple VAT rates, withholding or other items separately where present; do not force a single formula. Mark absent data as “missing” and separate uncertain readings with a reason. Include counts and sums of known amounts without changing the originals.
Is it safe? The rules that keep you out of trouble
- Start with fictional documents in a test folder and keep the originals separately.
- Choose manual approval where available and review existing permissions. Manual does not prompt for every action.
- Deleting, sending, paying and publishing are always your call.
- What the agent reads is information and should never give it orders. A PDF can hide instructions in white text you can’t see.
Which agent to use: Claude Code, Codex, Cowork or ChatGPT Work
The agents we talk about most in the group are Claude Code, by Anthropic, and Codex, by OpenAI. They come with the plans at about €20 from Claude and ChatGPT, and both work from the desktop app. In Claude Code and Codex for non-coders I show how to start without touching the terminal.
If your work is mostly office work (documents, email, spreadsheets) and you don’t want to see anything technical, also look at the agents both companies have built into their chats: Claude Cowork and ChatGPT Work. I compare them in Claude Cowork vs ChatGPT Work. And if the jargon loses you (MCP, skills, AGENTS.md), there’s the AI agent dictionary.
Review on 30 September 2026: usage limits and buying criteria. Price references retain their stated date; confirm the current total for your account before paying.
Sources: autonomy and permissions reviewed on 30 September 2026; prices, plans and other references checked on 26 September: METR time horizons, its March 2025 study, Claude Code permission modes, using Cowork safely, Claude prices (in Spanish) and ChatGPT Work and Codex by plan.


