~/blog / working-with-ai-agents.md
What working with AI agents actually looks like: a real day of client work
TL;DR: Unless you’ve done it, you can’t quite picture what working with AI agents looks like day to day. So here’s one recent working day: a citation audit and correction across a client’s public web pages. I walk Claude through one page in Chrome and narrate what I want. It maps what every page needs, builds the runthrough, and a second agent audits the work before I give it my final read. The output is blocks of HTML, ready for me to paste into the client’s content management system – and after the pages go live, Claude opens them in my Chrome browser and checks that every citation leads to the right paper, so I can also have a look for myself. All of it runs as one session among, on this particular day, six or seven open sessions on my dashboard. I move between them when an agent finishes its task and it’s my turn to review.
The part nobody shows you
Everybody posting about AI agents and AI operating systems shows you a demo of their flashy system. But it’s not a real demo. It’s fake, it looks pretty, and it’s showing you bells and whistles. Thirty seconds, one prompt, confetti. I have yet to see a real working day – the actual shape of sitting down with client work and a deadline.
I know why that is. A day like that is really hard to film and document while you’re building it and getting your work done, all at the same time. That’s why there aren’t screenshots from this particular day: I didn’t make any. I was busy building, working, and noting what worked and what didn’t.
But for someone who has never used a system like this, it is impossible to picture and difficult to imagine – and it has fundamentally changed how I work, and how I will work in the future. So here’s a snapshot of the other day, told exactly how it went. (More or less.)
The job on the table
A client’s website cites scientific papers – or it should. Some pages have citation lists with inline markers but no hyperlinks. Some have the lists and markers but no URLs anywhere, so the references have to be traced first. Some pages carry a pile of loosely related references and “further reading” instead, and the right spot for each inline citation has to be worked out backwards from the reference itself.
Different states, one goal: every page ends up with clean inline citations that link to the right papers. Period.
The final delivery: blocks of HTML I can paste into their content management system to update the pages.
That’s the job. It’s grunt work for someone doing the kind of work I do – an inevitable part of it, the kind of thing an intern could do. But there is no intern, and the time it would take to explain it to one, have them do it, and check the result myself doesn’t make sense. It’s the kind of job that can take hours when I do it alone, manually, page after page, reference by reference.
One thing I should say before I get into it: everything Claude reads in this workflow is the client’s public, already-published website – pages anyone can open in a browser. Drafts and internal material don’t enter it. And this client knows I’m using an AI-assisted system on their project. The full set of rules I work under is its own post: how I keep proprietary information out of chatbots.
Showing, not prompting
Here’s the part that might be a complete surprise to you: I do not write a clever prompt. I open the page in Chrome, with Claude looking at it alongside me, and I talk.
This block of text stays. That’s an H2 header, keep it. That’s bold on purpose. All the existing formatting survives – you’re doing this one thing, nothing else.
It’s the same way I’d brief a new colleague, pointing at a screen we’re both looking at. And it lands better than any prompt I could have engineered, because Claude is looking at the real page, not my description of it.
And if I describe something wrong the first time and decide I want to say it differently, I just say: no, wait, scratch that, this is what I mean instead. There is no going back and scratching your typewritten instructions for your intern. And Claude doesn’t get confused by all of this.
A system grows out of the afternoon
From that one walkthrough, Claude goes wide: it assesses every page in the batch. What state is this one in? Which pieces does it already have, which are missing? Out comes a map – who has what, so who needs what done.
Then we build the runthrough for one page, and it becomes the pattern for the rest. And because I don’t trust any single pass – mine included – a second agent comes in behind the first and audits the work: do the citations actually match the claims? Did the formatting survive? Is anything missing?
And now might be the appropriate time to say: the Claude doing the work is an agent. The second agent auditing it is another agent, spawned by the first. It’s not a hard-defined agent I took the time to set up somewhere – I told Claude, “bring in a second agent to audit the work,” and it does.
My turn, its turn
Once the pattern holds, I tell it to run, and the work happens in the background while I do something else entirely – something else that looks exactly like this pattern, but for a completely different topic. When a batch is ready, it’s my turn – more on that below. I review the HTML blocks, fix what needs fixing, upload them, and set the pages live.
Then the closing check. Claude opens the live URLs – the actual published pages – in my Chrome browser and confirms the result end to end: does every citation open the right paper? Anything that doesn’t match gets flagged for me. And I have another look for myself: all the windows are open in my browser, and I look through and check. The job is done when the live pages check out, and no earlier.
The dashboard, and the day around it
“How do I know it’s my turn?” you might ask. Well, the whole citation session is one tile on a dashboard Claude built for me – at the moment, a plain web page showing my open sessions, what each one is working on, and whether anything is waiting on me. When a session needs me, a new row appears: a description, and a link out to what I need to review and decide. When I’ve done my part, I tell whichever session is open, and the row disappears.
Because the citation work is only one of the things running. That day there were six or seven sessions open in parallel – different projects, different topics, all for the same client. My part, most of the time, is moving from session to session when it’s my turn: an orchestrator, and a reviewer with very strong opinions.
In the early days of my wiki, running sessions in parallel went badly enough that I thought I’d broken the wiki – that story is told in the system post, and it ends with Claude advising me not to run sessions in parallel. But that advice did not match my ambitions, so I asked Claude to write rules that would allow me to work in parallel. And once I got my dashboard, I could finally see it all. It’s funny how “don’t do that” is now the normal way I work.
The wish list
Working like this means you also generate ideas or problems at an unsustainable rate. Mid-task, I’ll notice: This step should be automated. That report should look different. Wouldn’t it be great if… If I chased each one of these wishes, I’d never finish anything – it’s too many things, and I know how easily I get confused when I split my attention.
So there’s one more session open, on a cheaper model, and its only job is to track my wish list. I dictate to it as things come up: here’s what I wish, here’s the situation that made me wish it. It logs everything. Later, when I have some free time, I sit down with a dedicated session, open the wish list, and we work out what’s actually worth building.
Built while using it
None of this system existed in today’s form before the work started, and there is no final form. I don’t step back from my work and announce: I’m going to work on building my system now. The dashboard, the runthroughs, the way sessions tell me it’s my turn: all of it gets improved and built while it’s being used, because the friction shows up during real work and that’s the right time to fix it.
The session-end habit does the remembering for me. When I wrap up a session, it harvests what happened. Every so often we hold a retrospective across recent days: what broke gets fixed, and what went really well gets written down as the standard way, so the next run can start from there instead of from old memory.
What a day like this is worth
I want to be straight about the state of this: my system is new, and it’s not perfect. Some of the communication between the sessions and me still needs tuning. There are definitely pieces missing I haven’t discovered yet.
But the citation job that eats hours and hours by hand becomes a day where the grinding parts run in the background and my hours go where they’re actually needed: judging, checking, deciding.
If you want the machinery underneath this – the files, the memory, the vault – the whole system is mapped in one post. If you want to try the smallest version of it, start tiny: give an AI lasting memory of your work, one folder, one afternoon. And if you’d rather build a setup like this next to your own client work than read about mine, Work with me is the door.
New here? I’m Julie – the homepage is the two-minute version of who I am and what this is about. Came with one specific worry, like an AI that forgets you or lies to you? The blog page is sorted by exactly those questions – start at yours.
The newsletter
If this was useful
The newsletter is where I send what I learn next – a letter every week or so on what I built, what broke, and what I’d tell you to try. No hype, ever.
Double opt-in · unsubscribe anytime · GDPR-compliant

I’m a scientist by training and a science writer by profession: chemistry and biology, 14 years at the lab bench, 8 peer-reviewed papers, and regulated biotech and pharma clients since 2011 – work where being wrong has consequences. For the last three years I’ve used AI on that real work, and here I document what actually happened: what worked, what broke, and what I’d tell you to try next. My best tip: if I can do it, you can do it.
