TL; DR: Reusing Agile Artifacts to Avoid Accumulating AI Debt
In this video from the 76th Hands-on Agile Meetup, I walk you through the A3 Delegation System and show how it helps avoid AI debt by borrowing artifacts and practices from Agile, such as the Definition of Done and Retrospectives. If you’d like to download the corresponding canvases (the artifacts of the A3 Delegation System), you can do so below.
You will get a full set of PDFs, along with the guide to the A3 Delegation System, so that you can run the system with your own teams. Enjoy the video and let me know whether you consider the A3 Delegation System useful.

📺 Watch the video now: How the A3 Delegation System Helps to Avoid AI Debt — Hands-on Agile Meetup 76.
🎓 🇬🇧 The A3 Delegation System Founding Workshop — September 28-29, 2026
Your team already delegates work to AI: reports, research, customer feedback analysis, stakeholder communication, or parts of operational workflows.
But can you answer these questions without improvising?
- What may AI decide, and what must remain a human decision?
- What does “good enough” mean for this particular work?
- Who verifies the result before somebody acts on it?
- Who checks whether the delegation still works after the model or workflow changes?
If those answers live in one person’s head, or nowhere, your problem is no longer prompting. You have a delegation problem.
The A3 Delegation System gives you a practical way to decide what AI may do, hand over the work clearly, define acceptable results, and inspect the delegation over time.
During two hands-on sessions, you will apply the system to a workflow. You will leave with a clear understanding of how to apply the A3 Delegation System to your workflows so that team members or stakeholders can understand, challenge, and continue your AI delegation work. Everything you learn is directly applicable to your situation the next day. The class is in English.
👉 Join the Workshop Now — $199: The A3 Delegation System Founding Workshop — September 28-29, 2026
🇩🇪 Zur deutschsprachigen Version des Artikels: Wie das A3-Delegationssystem dazu beiträgt, „KI-Schulden“ durch die Übernahme von agilen Artefakten zu vermeiden.
🗞 Shall I notify you about articles like this one? Awesome! You can sign up here for the ‘Food for Agile Thought’ newsletter and join 35,000-plus subscribers.
🎓 Join Stefan in one of his upcoming training classes!
What the Video on the A3 Delegation System Addresses
- AI Debt and the Delegation Lifecycle: The delegation, documentation, and ownership failures you already know from Agile reappear the moment an AI workflow runs unattended and undocumented, and the six-stage Delegation Lifecycle gives you a place to catch each one before it reaches a customer.
- The A3 Framework, Assist, Automate, Avoid: Every task lands in one of three boxes, Assist (AI drafts, you own the outcome), Automate (AI executes on explicit rules with an audit cadence), or Avoid (work that stays human because failure damages trust), and the real risk is Assist quietly drifting into Automate because the output has looked fine for four weeks.
- Model Routing by Cost Tier: Good enough is a deliberate decision with an owner, not a default, so you match each task class to local, mid-tier hosted, or frontier models by cost and stakes, and you keep proprietary process data out of the big labs because that data is your actual moat.
- The AI Definition of Done per Task Class: One Definition of Done for everything works in Scrum, but breaks under AI, so you write a separate one for each task class across four levels: verification, provenance disclosure, data hygiene, and sufficiency tier.
- The Delegation Audit and Governance Trail: Drift happens slowly, and approving a draft in four seconds is rubber-stamping, not reviewing, so you run a monthly audit for output drift, sufficiency drift, reversibility, and review creep, and the artifacts you create along the way answer any CFO, CTO, or procurement due diligence question without extra reporting work.
iFrame
📺 Watch the video now: How the A3 Delegation System Helps to Avoid AI Debt — Hands-on Agile Meetup 76.
Transcript ‘A3 Delegation System’
The transcript has been slightly edited to improve readability. I apologize to all native speakers in advance:
So, ladies and gentlemen, welcome to the 76th Hands-on Agile Meetup. And how could it be different? We are talking about AI. And more importantly, we talk about AI and what can go wrong if you actually embrace it and really use it in your personal work, and particularly if you do this at a team level or organizational level.
Because while doing this for yourself is relatively trivial, right? We’ve all been playing around with these chat tools. But if you take it a step upwards, so to speak, and if you start using this at a team level, organization level, things get really tricky rather quickly. And I’m not just talking about exploding AI token bills, right?
There are a lot of other issues that you will immediately run into, and I collected a few of those. If you are familiar with Agile ways of working, you will immediately understand this is somewhat familiar in the sense that these are the same patterns that you’ve been running into for years, and guess what? They’re popping up again. So, AI debt, that’s a catch-all phrase from my perspective, so I use it to point to everything that basically goes wrong, can go wrong, can go in the wrong direction, and that creates some sort of issues that you need to remedy if you want to be successful in the long run.
And there are many aspects of those. So, for example, in my online course, one of the stories I use is that one of the senior engineers was responsible for creating an agent-based workflow that was doing some important stuff in the background, but was never documented, right? And the other people just noticed, okay, it’s working, and they were fine with it.
But once he left, no one could say why this was the case, right? So it just grew this way. Organically grown over time. And it might have dire consequences.
So our little play here really put a lot of stress on the sales team for the company, because an important prospective customer was really aggravated about this. You may have had several versions of how this might appear in real life already. Who decided, for example, that the status report you send out every Friday to your stakeholders runs unattended? Who decided that there’s no need for a human in the loop, so to speak?
Who defined what good output looks like? And who checked last month whether it still does? That’s the point, right? You really run into a lot of these kinds of trivial questions if you just have a superficial view.
However, it goes rather deep. So what I thought about is the different steps of how you may address this topic of how to make your decisions regarding AI transparent, how can you document them, how can you audit them, how can you help other people understand why you did it this way and not in another way, how you can package all of that to help other people understand what you’re doing, particularly your management level or your leadership level and probably your CFO, and when can you figure out that you’re probably on the wrong path, that something drifted in the background no one has noticed. How might that look like? And I ended up with this kind of Delegation Lifecycle, so that is comprising of six different stages.
And all of these stages, we’re not going to address all of them, but a few of those. And it always starts with stage one. What should AI do, or I prefer to start even earlier, should AI do this at all? Right, so this is the A3 framework. I will show you in a second. It should be a no-brainer, but often it’s not the case.
So the time when you could throw the most expensive frontier model at whatever problem you had is coming to an end. If you’re an individual and you still have these Pro, Plus, or Max subscriptions with the two big companies, you’re still getting subsidized at an enormous level. If you were to pay all of your usage at the API token level cost, it would probably be much more expensive to actually do this and achieve something here. That’s a luxury that companies no longer enjoy.
They all have to pay by API. And you may have heard that, particularly if you’re encouraging the use of AI, particularly in software development, that your AI token bills go through the roof relatively quickly. Uber mentioned that they blew through the complete 2026 budget for tokens within the first four months. And this can be really, really expensive.
Cannot see the form? Please click here.
Transcript ‘A3 Delegation System,’ Part II
So you need to think about, okay, what can I hand to what kind of model? Because, spoiler alert, you do not need Fable 5 to turn a markdown text into some basic HTML. It’s absolutely a complete waste of money, right? And it’s taking more time, by the way, too, right?
So, handover. There’s one thing to decide: what you’re going to do and what AI is capable of supporting you. But there’s something completely different if you make the decision focusing on a specific task. Because you actually have to document a lot of things.
What goes in there, what we were thinking about when we decided this, etc., etc. Next, of course, we have a Definition of Done in Agile. So, shouldn’t we have one in AI too? Particularly given that it’s a probabilistic technology. And depending on the context and what model you use, the output may vary quite significantly.
So we certainly should have a good understanding of what ‘Done’ actually means. So once we understand what done means, at least from our perspective, shouldn’t we then also inspect it from time to time, whether what would get delivered is actually meeting our Definition of Done, right? And finally, the last stage, roll this up, package this. Help the CFO understand the return on token spend.
Why are we spending all of this, and what do we get in return? Because at the moment, if I check various analyses, by the way, check out this weekend’s newsletter, there are going to be some great articles in there about the state of the economy, macroeconomics, and analysis and reports, really fun stuff, particularly one from Google that’s really good, where they try to understand, okay, where is AI actually helpful? What is AI doing?
What are the costs? Where’s the return on investment? And if you’re following a bit of this AI bubble talk, you may have recognized that given the capital expenditures, or CAPEX, that the big hyperscalers are throwing at the problem, some people really start to get, I wouldn’t say cold feet, they really start to consider… Okay, so these are the six stages.
And by the way, all of those have artifacts. So little canvases where you can put stickies. I have an Agile background, as you know, probably. So we start with stickies before we do something else.
But two things are not in this life cycle. On the one hand, it’s a task inventory. So think of it as your backlog of AI workflows. Very few companies I know actually do this.
Everyone starts working somewhere, and everyone is doing something, but no one has an understanding, okay, where are we doing what? And the other thing is team norms. You already made all of the decisions. You already decided, okay, I use AI for this and for that, and we handle data in this way, and that data, it doesn’t, my goodness, never put this into any kind of model.
This one can only be used anonymized. So, you already made a lot of decisions. So, why don’t you put them together and create your AI working agreement, which makes transparent what you decided in the past and actually helps other people understand what you’re doing here.
Just think about the classic use case of onboarding a new team member, right? So, A3, the Assist, Automate, and Avoid framework, is in three simple boxes. As you can see, they’re really trivial. So, category one, Assist: The AI drafts. You decide and own the outcome. It’s your judgment in real life here. So, it could be the draft of a Product Backlog item. Maybe you’re using Gherkin, and you ask a model to come up with a reasonable description of how this is supposed to work. Automate. AI executes on explicit rules, and there is an audit cadence built into it. Okay, so, for example, at the end of the week, you collect the state of your project from your project management tool, whatever you use, and you aggregate a report for your stakeholders, and you send this out Friday afternoon.
So, probably that will work on autopilot most of the time. Maybe you have a look at this before it’s sent out, but it’s more kind of, yeah, it looks good. And then you say, go. Probably it’s okay, right?
But you certainly need to check it more thoroughly from time to time, because there is always drift. Have you noticed how your prompts and your skills behave differently when the model underneath changes? It shouldn’t be a surprise to anyone, right? But guess what?
Very few people do this, actually. They run a check on those. They run an audit on those. They’re trying to figure out, hey, moving from Opus 4.7 to Opus 4.8, is there any drift we can see here?
Is the outcome changing? And if so, in what dimension? Can you run a few tests here, dear model? So then you’re quickly in the realm of harnesses and evals.
And you quickly now probably understand why companies that are heavily using AI, particularly in software development, are actually hiring more people than they make redundant. Because all of the systems need to be actually maintained, set up, and understood. So it’s not all that easy. And then category three, Avoid.
Work that stays human because failure damages trust. So, in a manager-reportee relationship, for example, you probably do not want to do this by email. And your model is setting up an email, and you run your feedback skill on top of it, and then you send it out, because that might probably not be a good idea. What else falls in the Avoid category? For example, you may come to the conclusion that you never put names into the model. You’re interested in patterns, and pattern recognition actually works without the names of individual team members.
So that’s the background of the whole thing, right? So whatever you decide, it falls into one of the three categories. And the problem that I see in real life is that sometimes things that are supposed to be Assist drift into Automate because it’s more convenient. Because, yeah, it looks good. It’s been looking good for four weeks. Okay, why would we check now? So, delegate execution, never delegate your responsibility. It always comes back to you.
No one’s pointing fingers at the model and saying, ” Bad model, bad model.” It’s always you who’s carrying the brunt of the feedback here. Routing. Stage 1 of the life cycle decides if and how AI touches the task. That’s the A3 framework we just had a look at. Stage 2, Routing, asks which model tier is good enough for this task class. Good enough is a decision, and that decision should have an owner. And it should not be some sort of default that somehow drifted in over time, you know.
Yeah, you know, I’ve seen 26 of these emails, and I said, by now, they look good, right? And we can use them. That’s fine. Well, it’s probably not the correct approach, right?
You should be putting a bit more deliberate work into this. We have three tiers, mainly classified by cost. You can run local systems, open source systems. They probably have near-zero cost per token because you can run it yourself, unless you want to have one of these quasi-open-weight frontier models like Kimi or something like that.
This requires some serious hardware to make this work, right? But there are a lot of much smaller models available around that you can easily run on a Mac mini, for example, and then it’s basically free, right? Runs on your own hardware, high volume, low stakes work, or data that cannot leave the building. The latter one is important.
Even if you anonymize data, you probably do not want to hand your data over to one of the big frontier labs, right? That’s the paradox of using AI. You’re not only paying for the tokens, but you also pay with your data if you send them over to ChatGPT or whoever you choose to work with, Gemini, or whatever, right? And if there is one moat, it’s your data, particularly your data or process data that you as a company collected over decades.
That is the interesting part. Siemens is not doomed because it stopped building refrigerators. No. They are the world market leader in industrial automation.
They have terabytes of data on how to run operations at an industrial scale. That is the interesting part. That’s what everyone is after. There’s a reason why Nvidia has a pilot project with Siemens.
That’s the interesting part, and you want to keep that data for yourself. Mid-tier hosted, moderate cost, enough unstructured input drives the output, and the task needs no frontier reasoning. You have a chatbot that is doing first-level support. You run Fin, for example, to help your customer care agents to focus on the real work and cover the 30-40% that you have at the beginning by something that is actually not bad.
And seriously, I have several software providers where the first-level support agent is sometimes better than a human, because they have access to all the documentation. They can exactly say what… If you ask, okay, where is what? They probably have the answer.
That’s really good. And then, of course, tier three, the frontier model tier, high cost, only when the task needs the strongest reasoning, and the output value covers the bill. So, software development is a classic use case. So this is something that you need to figure out.
There’s more and more software available to do this. Ramp has just created a router for this purpose, Cursor 2. So it’s not that you have to start from scratch, but again, you need to do this. And actually, you can do this at the Post-it level if you want to.
You just need to sit down, run a few experiments, and then you have a rough understanding of where this whole thing is heading, what model might be appropriate. It’s not that you have to throw technology at this from day one, but you have to think about what you’re doing here. So, the Handoff Canvas.
Generally, deciding what AI is allowed to do is great, and to what extent it’s allowed to do. But for every recurring AI workflow, you should actually go through these six questions. Task split. What does the AI do? What does the human do? And where is the boundary? And this is floating for every single workflow. No two are identical.
Inputs. What data does AI need? In what format? What must be anonymized?
If you want to have a really, really good working experience with AI, it’s less the question of throwing the latest model at it. It’s a question of having the right context. And the right context is also largely influenced by the available data. If the data is clean and structured, you would be surprised what older or smaller models can deliver.
Garbage in, garbage out. Yes, these models are great at going through large amounts of text, for example. But they struggle to find the needle in the haystack, similarly to how we do. Outputs.
What does good look like? Format, length, and quality criteria. This is actually the box for the AI Definition of Done. So we will have a closer look at that one.
Validation. Who checks the output? How does it work? Is it an agent doing that? Are you running some evals in the background?
And only if the evals fail, 15% of the evals do not work, then please get in touch with your human. Could be. Against what standard? Yeah, hopefully that’s your AI Definition of Done, right?
Using what method? Again, is it human doing? It’s good, I know it when I see it, the classic approach, which is basically really hard to document and help other people understand because it’s somehow part of a human.
Failure response. What happens when the output is wrong? What are the stop rules? And who owns escalation?
Does it mean you pull the string and the assembly line stops? Or is it, who cares? Well, happened in the past. Don’t worry too much, the output is a bit off here.
It’s drifting. No one cares. Just send it out. You need to make a decision on that.
Records. What do you log? At what level of detail? And who owns the log?
And what do you do with that? And as you can see, we’re just halfway in. It’s getting more and more interesting, right? And it actually sounds quite reasonable.
This is not the playbook of an AI doomsday. You’re trying to avoid the whole thing, right? This is not what it’s about. It’s about the professional approach to dealing with these issues.
So, the Definition of Done has four dimensions per task class. So it’s really important. In Agile, we always pride ourselves on having one Definition of Done for everything, for whatever we do. This does not work for AI, absolutely not.
So what you need to come up with is an understanding, starting with the A3 framework, of what a task class actually is. The AI Definition of Done needs to be different for something where you have Assist compared to something that is Automate. These are two different things, and you can’t throw them into one bucket, so to speak, because you’re lazy and you just want to have one. My suggestion is to create one for each task class and then consider four DoD levels.
So, verification level, what gets checked by whom and how, and what is the source? Provenance disclosure: Which label applies, Assist or Avoid? It would probably be a good idea to have an understanding. Okay, this task class is about Assist.
It’s not Automate. If we start to have it drift into Automate, something is wrong. Data hygiene. We already mentioned that.
What never enters a model? Very simple. Sufficiency tier, coming back to Routing, which tier is enough, and why? So you see, they’re all intertwined.
That’s the purpose. Because it makes life so much easier. So, of course, I created these little canvases here where you basically have this one example. By the way, I’ll send you the PDFs later.
Where you have the examples, and you can check what this is actually about. Verification level. Every claim about feature status is checked against the release notes by the sending manager before sending, every single time. So in this case, they agreed on, okay, before we send the status report to our stakeholders, we check the draft against the release notes.
Because it is so important to us that this is on the right track, and that we’re not promising something to other people that actually is not true. Once bitten, twice shy. So, delegation audit, what do you do? It’s one thing to define a standard, and that you all agree, yes, nodding, yeah, we need to check this from time to time.
But when do you do this? Of course, in Agile, we have the Retrospective. So, probably it’s a good idea to have the same stuff here, right? So what are classic checks that you should have for each workflow in a regular cadence?
Output drift. Outputs checked against the AI Definition of Done. Very simple. Are our outputs that we generated over the last four weeks still acceptable if we compare them to the AI Definition of Done?
And yes, of course, you can create some agents that do this on the side. And probably it’s a good idea to use a different model that’s doing that, et cetera. So you can automate this to a certain extent to have some built-in triggers and alarm bells. No problem.
But sooner or later, there will be a moment when you have to actually be the human in the loop and say, ” Okay, I have a look.” Because otherwise you create a self-referential system that’s not really helpful from a governance perspective, right? Sufficiency drift. Is the assigned model tier still right?
Is it still working? Or should we probably consider a different model tier or a different model? Sometimes you just need to figure this out. And in my experience, drift happens slowly.
And you probably won’t even start noticing it at the beginning. Reversibility. Could you stop this whole thing today? Could you pull the cord and say, stop our automation?
Our AI-supported workflow doesn’t work anymore. We need to check this before we can continue working with this. I thought about this for quite some time. Creep check.
Check the review time. If you are the human in the loop and your job is to do some reviews, and you approve a draft in four seconds, this is not reviewing. This is rubber-stamping, right? And you don’t want that.
If that happens, you’re probably already in deep trouble. By all means. That typically means that you waited too long to do this. This is also the reason that I suggest that you actually do this on a regular cadence, every single month.
It’s like having a Retrospective. I don’t know, the last Thursday of every month, you run your checks here. Okay, and of course, I created a few little cards here. So, output drift, pull three recent outputs, check line by line against the Definition of Done.
And then you say, okay, pass, drift, fail. And the same with all the others. And then you’re actually on a good track. It’s not about being right all the time. It’s about catching drift before it creates damage. And damage comes in a lot of different shapes, sizes, and colors. So what else?
Due diligence questions. So, wrapping all the things up so that you can actually talk to other people. The cool thing about this system is that if you go through all the artifacts, if you support the whole process as it is supposed to be, you actually create more than enough artifacts and documentation that you have a rather relaxed approach regarding any form of audit or governance inquiry or whatnot.
So if a prospective customer asks you how do you govern your own internal use, you actually have the documents already available. You have the A3 decisions from the delegation portfolio. So your CTO or your CPO will really be excited about this. You have routing records that actually show what you spend on each task class.
Your CFO will be really, really pleased about that. And you also have sign-off and audit logs from the audit trail. So if you need to talk to the procurement department of your customer, you can show what you’re doing. The records you have, all the artifacts, actually do the reporting.
That’s also the reason why this whole system does not include some sort of leadership canvas or something like that, a leadership dashboard. We have it all already. All we have to do is probably create a skill or an agent that is actually collecting all of this information into a dashboard. And that shouldn’t be that much of a problem.
So my suggestion, if you’d like to try this at home, pick one workflow your team already runs today. Walk through the six stages and start your inspection by writing down the first stage where you have no answer. Because when you have no answer, it’s actually a good point to start. And start where your pain is, and for many teams this will be stage four, you know, that’s the Definition of Done, because I don’t know a lot of teams that actually went through the hassle of defining what done looks like when they’re working with AI.
Yeah, and that’s it from my point of view. So now, probably these six steps make more sense. Decide what AI should actually do, which model is actually useful to do that. Okay, what needs to be decided?
How is the work transferred to the AI? And you do this for every single workflow. Define a Definition of Done for every single task class where AI is involved. Regularly inspect what you are doing here, so basically have a sort of Retrospective or an audit. And actually, in the end, you can roll everything up into a nice package for the leadership, because you already have created all the other stuff that you need.
And two things that are not here. So maybe here you have the repository of where you use AI for what purpose. Remember, you also need to come up with an idea of what a task class is, because otherwise your AI Definition of Done is not working.
And on the other side, you basically have your AI working agreement with your team. Okay, how are we doing this? When are we doing this? Who is responsible for what?
For example, it has really proven to be helpful if important decisions have an owner. If an artifact has an owner, for example, it makes life so much easier if there’s one person who can explain why and how. Okay, that’s it from my side. Thank you very much for your attention and patience.
A3 Delegation System Workshop — Related Articles
You Already Have an AI Working Agreement. Write It Down.
If You Can Facilitate a Retrospective, You Can Audit Your AI
The AI Definition of Done: Human in the Loop Is Not a Quality Standard
The AI Delegation Lifecycle: Your Team Has AI Outputs. Where Are the Decisions?
Assist, Automate, Avoid: How Agile Practitioners Stay Irreplaceable with the A3 Framework
The A3 Handoff Canvas: Six Questions That Turn AI Delegation Into a Repeatable Workflow
The A3 Framework: Assist, Automate, Avoid — A Decision System for AI Delegation
No More Cheap Claude: Four First Principles of Token Economics in 2026
👆 Stefan Wolpers: The Scrum Anti-Patterns Guide (Amazon advertisement.)
📅 Training Classes, Workshops, and Events
Learn more about AI Builders with our AI and Scrum training classes, workshops, and events. You can secure your seat directly by following the corresponding link in the table below:
| Date | Class and Language | City | Price |
|---|---|---|---|
| 🖥 💯 🇬🇧 July 20,2026 | GUARANTEED: AI 4 Agile Course v3 — Master AI for Agile Practitioners (English; Self-paced Online Course) | Self-paced Online Course | $149 incl. 19% VAT (If applicable.) (Before and after: $249) |
| 🖥 💯 🇬🇧 July 23,2026 | GUARANTEED: HoA 76: Where Are Your Decisions to Avoid AI Debt? (English) | Meetup | FREE |
| 🖥 🇩🇪 August 25-26, 2026 | Professional Scrum Product Owner Training (PSPO I; German; Live Virtual Class) | Live Virtual Class | €999 incl. 19% VAT (If applicable.) |
| 🖥 💯 🇬🇧 August 27 to September 17,2026 | GUARANTEED: AI4Agile BootCamp #8 (English; Live Virtual Cohort) | Live Virtual Cohort | €499 incl. 19% VAT (If applicable.) |
| 🖥 💯 🇬🇧 Sep 28-29,2026 | GUARANTEED: A3 Delegation System Founding Workshop (English; Live Virtual Class) | Live Virtual Class | $199 incl. 19% VAT (If applicable.) |
| 🖥 🇩🇪 Sep 30 to Oct 1, 2026 | Professional Scrum Product Owner Training (PSPO I; German; Live Virtual Class) | Live Virtual Class | €999 incl. 19% VAT (If applicable.) |
See all upcoming classes here.
You can book your seat for the training directly by following the corresponding links to the ticket shop. If your organization’s procurement process requires a different purchasing approach, please contact Berlin Product People GmbH directly.
✋ Do Not Miss Out and Learn More about the A3 Delegation System Workshop — Join the 20,000-plus Strong ‘Hands-on Agile’ Slack Community
I invite you to join the “Hands-on Agile” Slack Community and enjoy the benefits of a fast-growing, vibrant community of agile practitioners from around the world.
If you would like to join all you have to do now is provide your credentials via this Google form, and I will sign you up. By the way, it’s free.
The post How the A3 Delegation System Helps to Avoid AI Debt Borrowing from Agile Artifacts appeared first on Age-of-Product.com.







