Inside AI - Ep 5: Who Answers for the Agent
In this episode of Inside AI, Sam Bradon is joined by Ritwik Singh, Founder of 36ARC, to unpack who actually answers when agentic AI makes the decisions.
Ritwik traces agentic AI through three phases, from proving it could work, to proving it does work reliably, to the current question of accountability. He argues capability and value have largely been answered, but governance has not kept pace.
The conversation turns to financial services, where APRA's recent supervisory letter and a landmark lecture by NSW Chief Justice Andrew Bell both point to a widening accountability gap. Ritwik notes that deferring to a confident AI system doesn't discharge a director's personal duty of care.
On operationalising governance, Ritwik offers three anchors: keeping deterministic control over where probabilistic judgment is allowed, building traceability into every decision, and naming a person who answers when a system gets it wrong.
A sharp, practical discussion on why trust in AI is an operating model, not a product feature.
Runtime [00:35:35]
Our Speakers
Sam Bradon
Sam Bradon is a seasoned professional services leader with deep expertise across sales, delivery, and customer success. With a track record of leading large consulting teams and solving complex client problems, he helps CTO Consulting clients design and deliver innovative, strategic, and impactful business solutions.
Ritwik Singh
Ritwik Singh, Founder 36ARC, is a seasoned technology leader with deep expertise across digital transformation, automation, and enterprise strategy. With a track record spanning banking, government, and global consultancies, he helps organisations design and deliver AI-powered, customer-centred process solutions.
-
Speakers: Sam Bradon (Host), Ritwik Singh (Guest)
Sam Bradon: Hello, and welcome to another episode of AI with Intent. This is a podcast where we focus on providing unfiltered and real-world views on our experience of AI in real life, based on our real-world experience. I'm Sam Bradon, I'm your host today. I've been in IT now for about 30 years, both in the UK and Australia, and over that time I have had a real focus on helping organisations get real value from their investments in IT.
I am absolutely delighted to have Ritwik Singh with me today. I've had the absolute pleasure of working with Ritwik now for nearly 15 years, and a fair few gigs. He's someone I've always enjoyed working with, because I love the way he tries to understand what the real problem is that we're trying to solve.
Ritwik is the founder of 36ARC, a Sydney-based advisory that helps organisations cut through a lot of the noise around AI and automation to get to the real value. Ritwik has over 20 years' experience in IT, including over 15 years of Pega experience, and amongst the many clients he's worked with are CBA, NAB, Optus, and federal government, as well as very successfully scaling out a Pega partner for many years. Ritwik really focuses on working with enterprise leaders, helping them separate real operational value from some of the AI hype.
So, Ritwik, welcome.
Ritwik Singh: Thanks, Sam. Great to be here, and a very topical thing we're going to be chatting about today.
Sam Bradon: Well, AI is getting a little bit of noise nowadays, isn't it? Just a little bit. I should just say, before we get started, that the views we're expressing here are our own personal ones, rather than those of the organisations we work for or with.
So, moving on to today's topic, and to set the scene: Agentic AI has really moved from the conversation around its potential — doing pilots and proofs of concept — through a number of proof points around whether it does or doesn't actually work, through to more about who's actually accountable for the decisions and outcomes AI makes. So the narrative is really turning towards governance and explainable AI. Today's question is: who answers for the agents?
I'd like to start by asking, Ritwik — where do you think this conversation actually is, as opposed to perhaps where the marketing is leading people to believe it is? I think Agentic AI has moved quite quickly through a number of distinct phases in an exceptionally short space of time, and perhaps some of the published material is lacking on where the technology actually is at this moment. So where do you think we are right now?
Ritwik Singh: Yeah, no, absolutely — with Agentic AI, everything seems to be agentic. I'd really put it as three phases. And I think you're right that the writing and the narrative behind it really lags by a fair margin in terms of where we're at and what's being said about it.
I think the first phase was all about potential, right? We all got excited, but we also wanted to know, in an enterprise setting: could it even do the thing? Could it really achieve what we set out to achieve? It was all about demoing the capability — and some pretty impressive demos, where you could see, "look at this thing, you can draft a letter, draft your emails, summarise documents."
Then it got a little more nuanced. I remember there were pilots everywhere, proofs of concept, a lot of people talking about the impact it was going to have, and a lot of wonder. Based on all of our collective personal experience, I'd say that's fairly earned, because it's really arrived a lot faster than most of us expected. I don't think we've seen things happen this quickly.
But then the second phase — this is where the rubber really started hitting the road, and a lot of organisations have been living this over the last year, year and a half. It's a much harder question they're trying to answer: can we actually get this thing to work? Not just in a demo — a demo using a very supportive environment to make it work — but in real-life operation, everyday situations with real customers, real data, and, more importantly, real consequences.
So "how reliable is it, really" becomes the question. That's turned out to be a much harder question to answer, because it's not really a question about the models — I think the models are heading towards that level of utility. It's starting to become more about, "can I rely on this?" And that's when we get to the third phase — this is where we have agentic systems, or agentic AI, being unleashed to start making decisions. Then we've got the question of who is actually accountable for those decisions, and that's a question everyone is walking towards. Some people — especially in marketing — do have some answers for it.
The first two were all about, "can this thing do it?" — capability — and then the proof of value: can it, and does it? I think we've answered those. But the third becomes about governance. It's more of a legal question, and in some cases a question about a specific person's liability — a personal liability, right?
So we've moved away from asking what the technology can do, towards asking who actually answers for it. And I think they're different questions.
Sam Bradon: Yeah, look, I'd agree. And perhaps, before we talk about that third phase, maybe just on the second one — we've all seen hundreds of demos from that first phase, and we're starting to see some in the second phase, but I still feel organisations are getting a little stuck there. It's perhaps proved slightly harder to turn a demo into something enterprise-ready. Why do you think that might be?
Ritwik Singh: Oh, look, I think there are a lot of factors, Sam. The old cliché is that every organisation is going to be different — different context, but that's the reality: all the environments are different, their stack might be different, there are varying levels of complexity, especially around the type of use cases you try to attack with the technology. And then there's the terminology we all love to use, especially at the start of something — and technical debt.
But the other big one, I think, is data — what's the state of your data? How usable is your data, and how mature are you as an organisation to push pretty disruptive, groundbreaking technology through? And then there's the whole picture of change management effects as well.
This is where — and I'll come back to that word again — that issue of reliability comes back. The things you're trying to solve have to do the right thing reliably, every time, regardless of how messy the inputs get. Then you've got all the edge cases starting to come through — scenarios you didn't anticipate — and that's where it starts to show.
So a lot of organisations have found — and we'll leave the token economics conversation to one side for now — that you've spent all that money, and can the agent do it on its own? I don't think that's enough to get them there. It needs structure. It needs to know — the agent, in this case — what it's allowed to do and what it's not allowed to do, so that when it hands over to a person, it can give some context. It needs to know when to hand off to a person, and what happens when it's unsure — what the steps are, in what order, with a record of who did what and why.
Sam Bradon: So how do you make sure it doesn't hallucinate, and actually tries to give you the right answer?
Ritwik Singh: Yeah, and hallucination came up pretty early in the piece when the Gen AI movement started. There are lots of different technical terms floating around — RAG, retrieval-augmented generation — giving it facts so it has more context. And then orchestration — for you and I, with our backgrounds, when we hear "orchestration" and "workflow" we probably envisage a different world, but across the board you'll see "workflow" and "orchestration" mentioned a lot.
I think if you boil it down, it's the structure that makes the difference between something that's very clever and looks good and shiny, and something you can actually reliably run your business on.
Sam Bradon: Yeah, and I think that's where you've got to have measurable value coming from something that's capable. I'd agree — and it probably leads us neatly onto your third phase, which is really around accountability. We've both worked a lot in financial services, which is regulated by APRA here in Australia, and they actually released, effectively, an open letter, didn't they — back in April, I think — providing a number of observations around the use of AI and the responsibility and accountability for it in financial services. I'd be interested in your take on what that letter was, why they sent it, and what the implications were.
Ritwik Singh: Yeah, it's pretty interesting for them to actually issue a letter. Regulation is always going to have a bit of a lag, but it helps to put it into context of why financial services first — particularly in this country, some of the big banks are quite ahead of the curve with this technology. But banking, insurance, super — everything APRA covers — they're the most heavily regulated, and they've been making automated, consequential decisions about people for decades, right? We've been involved in some of those transformations — everything from how you assess credit risk and price loans, to how you assess claims — that becomes a pretty significant moment when someone puts a claim through. And obviously a very big topic has been fraud and anti-money laundering, and everything around that.
So they're definitely further down the road, and some of the bigger players — some of the big banks — are actually quite operationally mature. But they're watched most closely, and if you want to see where the accountability question bites first, that's where you've got to look. This is where APRA comes into play — they went and did a targeted supervisory review, engaging with a whole bunch of the largest banks, insurers, and super trustees, and examined how AI is being used inside them. This was beyond just the polished presentations and the marketing spin — they started to look at the real risk frameworks being put in place, and what the accountability chains were.
Then, at the end of April, they wrote that letter — an industry letter, about what they found. They don't usually do that when things are going well. The headline was simple, and not surprising: adoption is fast, everyone's excited and wanting to do it, but the governance around it hasn't kept pace. That exposes institutions to cyber and operational risk. The cyber side gets a lot more airtime, especially with some of the models that came out and got pulled back, and everything that's happened around that.
But underneath it all, what actually compelled them to write it is an accountability gap that's becoming wider and wider — decisions are increasingly being made, or quietly shaped, by these systems. That's what APRA was really looking for: who's accountable for those decisions.
Sam Bradon: Yeah, and I think they particularly called out the role of boards, didn't they — around whether they were really stress-testing, or getting any independent auditing on some of the products, or the vendor promises around what was being included in the products.
Ritwik Singh: Yeah, absolutely. It's not new for APRA to have this level of scrutiny. In reading the letter, and preparing for this podcast, it took me back to when I worked on a pretty massive transformation at CBA a long time ago — 15 or so years ago, dating myself, probably just before we met, actually.
Sam Bradon: Yeah, yeah, I know the project you mean.
Ritwik Singh: Yeah. That was all about how the bank priced its big-end-of-town business and institutional loans — a lot of complexity, a lot of rules, a lot of tech debt, but also a lot of bespoke practices in how the models worked. We took something that was very slow and very IT-dependent, in terms of how the bank made changes to prices and assessed loans, and gave the business the capability to make real-time pricing decisions. But it was sitting on some pretty complex stuff — a complex credit risk calculation and portfolio simulations. There was AI in it — and nowadays we have to say "old-school AI," because there is a distinction — a Moody's risk engine running thousands of Monte Carlo simulations to produce a distribution of outcomes, and that's what was then used. It wasn't a deterministic calculator you could just plug into Excel — it was probabilistic, quite sophisticated.
But that model needed validation, and the process to get it validated, before it could be unleashed to the bankers, was very comprehensive. Alongside building the technology solution, we worked with the business to stress-test the model. Then it had to go through an independent audit, get signed off internally by the risk committee, and be presented back to the regulator before it went live. That was an ongoing thing — every time you make any material changes, you keep that process going. It was a very critical platform, and it got the scrutiny it deserved. But this is the catch — this is the risk running the other way now.
Sam Bradon: Correct. And, as you just mentioned, probabilistic decisioning is actually nothing new — we've been working in that space for 15 years or so, as you say. AI brings additional complexity, because we'll probably talk about explainability of decisions.
It was interesting, though, because I also noticed it wasn't just APRA — there was a whole run of other people or organisations speaking out about AI and governance. I think there was something about the Chief Justice of New South Wales also talking about it. If that's the case, that's quite important, isn't it? Do you want to expand on that?
Ritwik Singh: Yeah, absolutely. Timing-wise, APRA came out with a letter, ASIC had something as well that was more aligned to the cybersecurity side, and then there was this lecture by the Chief Justice of New South Wales.
It's very striking that the Honourable Andrew Bell, Chief Justice of New South Wales, delivered a lecture — the Harold Ford Memorial Lecture, not something I was completely aware of, I'll admit, until this topic resonated — but it's one of the most prestigious lectures in Australian commercial law. He used the occasion to speak on individual directors' duties in the era of AI.
I think that's a bit of a signal of how far this thing has travelled. The argument was about accountability being personal, in a way that APRA didn't talk about — APRA talked about how the institutions operate, how the banks operate — but his focus was on the individual director's duty of care, and what protections they rely on when executing that duty in making a decision. His point, in plain terms, was that the protection only holds if the judgment was actually yours.
It's the same principle, and it's standing up in real judgments too — there's a precedent being set. The Star Entertainment Group went through some civil proceedings, and Justice Lee of the Federal Court made the point directly: AI tools can help directors digest heaps of information — there's a lot of data to digest, and this applies across all levels — but it can't displace human judgment. That was the call. There's no substitute for management acting with integrity.
He cast AI as an amplifier — something that helps you comprehend information better, but can't replace independent thought. He was pointed about directors controlling the information they receive, and that they cannot use the sheer volume of board papers as a reason to disengage from it.
So it's very interesting — if a director simply adopts what an AI system recommended, they haven't exercised judgment, they've transcribed it. And the protection they'd otherwise have falls away. Chief Justice Bell said he'd want every board member in the country to hear this: that AI produces answers with great confidence and clarity of language — maybe a likeness to nefarious actors in our lives, fraudsters — that's what they sound like. So he's basically telling them: you can become fluent, and it's really important to be fluent, but that's not the same thing as being right. Deferring to a confident machine doesn't discharge you from the duty you personally carry.
Sam Bradon: Yeah, yeah, I certainly get that. I think one of the other things APRA mentioned, still talking about boards, is whether boards had the right skill sets to challenge some of the vendor presentations. We work for — well, we work for the same vendor, and other vendors — and everyone has an AI story, and has to at the moment, because that's what people want to buy. But it's almost an uncomfortable thing to say — do boards have the right capability to actually challenge some of the vendors?
Ritwik Singh: Yeah, no, absolutely. I think we have to be very careful here, because it's not about being critical of boards, per se, or vendors — that's not the intention. There are some vendors out there who will just slap AI on their presentations, but you and I, as you mentioned, have been on the vendor side, on the same platform, at the same time, and we know all about smoke and mirrors.
But if there's conviction behind it — you can make compelling demos, you do want to show the art of the possible, but you also want conviction behind what you're putting out there. I think the education needed for boards is about this technology, and that black-box aspect. When we were pitching to Rabobank back in the day — a very successful pitch — there was a lot of sitting with people and walking them through: what is this technology capable of, how is it actually going to impact how you do things differently? And it worked.
So it's a good thing if you're doing things with that level of conviction — and, obviously, we have a bit of a biased view that the sector needs more of it, not less. I think that's what APRA is really trying to point at — not that vendors are dishonest, but that there's something more subtle going on. The technology is very powerful, and it becomes about whether you, as the audience, are capable enough to understand what's going on, especially in the black-box parts — and whether you, as the vendor presenting this, might be doing it with all the good intentions, but shouldn't pretend to fully understand how that black box works.
Most of us have been using these tools personally for a while now, and we know we can tweak our prompts and do different things, but we're not necessarily always getting the same answers. That's how you have to treat these things — as self-learning, autonomous systems that are also trying to make decisions in this regulated way.
Sam Bradon: Yeah. So does that imply that things like a control plane, guardrails, AI control towers, are the answer here?
Ritwik Singh: Yeah, look, from a technology perspective it's definitely part of the answer, but we have to be careful, especially as people rush towards it — it's not the full solution, because it's very easy to say, "I've got something that oversees this, a control-tower type of concept, and I can give you an explanation of everything that's happened."
That explainability piece — it's easy to say, but with large language models, or LLMs, it's not necessarily a faithful window if it's reconstructing what it might have done. It might be a very plausible explanation after the fact, but I think it's more about traceability: what information did the system actually retrieve as part of that execution? What policy version did it use? You might start to see a proliferation of this sort of thing happening in your organisation as you get tools that help you work faster — but what is it also permitted to do? That's where the constraint comes in. As part of its execution, what tools did it use, and what recommendation did it make? Which deterministic checks were applied? The human-in-the-loop aspect comes into play here too — who approved it, who overrode it, was it allowed to proceed past a certain point?
And was the human in the loop just there as a tick-box exercise? There's a whole narrative coming out about humans in the loop having too much to do, leading to a bias towards just approving whatever comes across your desk, because you start trusting the system. So you need that evidentiary chain that can be interrogated by humans, so you can actually defend it — and in a regulated setup, that's very important.
Sam Bradon: Yeah, okay. So, really, governance — or governance tools — are very important, but only one part of it. The hardest part is almost non-technical, in some ways.
Ritwik Singh: Yeah, absolutely. That's where you've got to start as an organisation — deciding where you're actually prepared to delegate judgment, and what your system can absorb. Are you ring-fencing it, boundary-fencing it in a certain way? But you also want to get the benefits of the technology — you don't want it to become a lowest-common-denominator situation, where this doesn't look too different to systems we had a few years ago.
So — and this is particularly important for the agentic piece — what can it execute on its own, and under what conditions must it stop and escalate? That's where the almost-autonomous piece starts to kick in.
Sam Bradon: Yeah, so you've almost got a deterministic piece, which is like the control — would you say — and then there's the probabilistic piece. How would you expand on that?
Ritwik Singh: Yeah, that's what I was touching on before — I've seen stuff come out in the last three to six months where there's a narrative forming around, "I've got a way to control this AI." The agentic AI is doing things, and when you look at it, it's actually dwindling the benefits of the agentic piece, and starts to look very much like workflow and orchestration tools.
I think the objective is deterministic governance of where that probabilistic judgment is actually allowed. A rules engine can calculate, and does that robustly; a workflow engine can enforce all those mandatory steps; and a control plane can constrain those permissions. Let the LLMs do what they're good at — interpreting messy reality, handling variations, proposing or taking actions — but within a bounded operating envelope.
None of that removes the accountability. You can buy tooling that records what happens, systems that enforce boundaries, dashboards that surface risk — but you can't buy the person who answers for where that boundary was drawn in the right place.
Sam Bradon: Yeah, and where the resulting decision was one the institution should have made.
Ritwik Singh: Yeah.
Sam Bradon: And I suppose that leads to the whole — I think we've talked about this in other podcasts — but trust. What an organisation wants is to generate trust in its systems, and therefore in its customers, doesn't it? Trust isn't typically a product capability, is it?
Ritwik Singh: No, exactly. This is going to sound a bit clichéd, but trust is an operating model. It comes from being very explicit about that delegation, retaining that evidentiary chain, and having that evidence retained in a way that's non-destructible. But your escalations also have to be meaningful — you have to monitor your outcomes, and have a named person prepared to actually answer for when the system gets it wrong.
So the sooner leaders stop shopping for the shiniest product, and absorb that responsibility themselves, the sooner they'll start focusing on the real work.
Sam Bradon: Makes sense. So, in your opinion, are there organisations or places actually starting to adopt this control in a meaningful way?
Ritwik Singh: Yeah, absolutely. Some of the leading organisations, particularly some of the banks in Australia, are definitely at the cutting edge from a governance perspective. But taking an industry-wide perspective, the most detailed work I've seen has come out of Singapore so far.
That's worth noting — they're being very deliberate about not being anti-innovation with their approach. It's not that they're saying, "these are hard and fast rules you have to follow" — they're giving you a governance framework. In the last week or so, there's been stuff out of the UK on a sovereign AI model capability that banks are looking to sign up to, and India has talked about how they're going to approach accountability and transparency with AI — but that seems more like mandating rules.
The Singapore approach does a few things that make it work really well. The first is: don't try to supervise every decision. This comes down to why you're trying to automate in the first place, and not defeating that reason. You don't want a human in the loop at every point of your process — as I mentioned before, when a human in the loop suddenly gets a huge volume of stuff, they just become rubber-stampers, developing an automation bias, relying on the machine and just stamping it. That's not effective.
Instead, you identify the moments that really matter — the significant decision points — and have a human approve before the system acts on those points. This is all about establishing a tiering of your autonomy scale — a risk tier, for what's high-risk and significant versus low-risk. That's part of their framework.
The second move from the Singapore framework is that they name the thing that quietly underlies all of this — that automation bias, the human tendency to over-trust a system, precisely because it's been reliable in the past. The better these things get, the more we defer to them and trust them — and that's exactly how judgment leaks out of the process without anyone ever deciding it should.
The third thing — and this one probably matters most — is that their framework states plainly that governing these systems requires clear lines of responsibility. That comes back to what we were just talking about: you want human oversight across the whole process, not just as a kill switch, not just safety features built into the model itself — you want the organisation set up to answer for what the tool does.
So it's a framework they're proposing, not law, and they keep refining it — they've released some new material very recently. But they're trying to stick to their goal of not inhibiting innovation, while still giving an honest attempt to answer the question of how, rather than just insisting that someone should.
Sam Bradon: Yeah. Okay, I'm slightly conscious of time — we're getting towards the end of our slot. So the question is — for you, Ritwik — if I'm an enterprise leader, what would I do next Monday morning? What should my responsibilities be, to move this forward? Funny enough, we're recording this on a Friday, so hopefully—
Ritwik Singh: (laughter) No, look, I think it's an old-fashioned take — it comes back to knowing your why, and your what, before you touch the how. What outcome are you actually trying to produce — and the why, stated in terms a person can be held to. That's the crux of it, because once you're clear on that, two requirements follow — the two I'd hold any agentic system to.
Number one: it has to be traceable — meaning you can reconstruct what it actually did, what it drew on, what it was permitted to do, which checks ran, and who signed off. Number two: it has to be reliable — meaning it gets there by a process you can repeat and defend, rather than something closer to instinct that happened to be right so far, particularly in those happy-path use cases.
If you get traceability and reliability going, sitting on top of a very clear outcome as to what you're trying to achieve, that's where the technology starts to become very useful. If you skip them, fast-forward them, or gloss over them, you end up with automated decisions no one can explain, and unpredictable things happening out there. That becomes a very dangerous path.
Sam Bradon: Yeah. (laughs) Well, listen, Ritwik, thank you very much, as always — really enjoyed that conversation. A couple of things for me — you touched on earlier, which I think is very important, is that deterministic governance of where probabilistic judgment should be allowed, and how that flows up to the board, both as an individual and a corporate responsibility.
So thank you very much for your time — always good to chat to you, my friend — and I'll buy you a beer soon. How's that sound?
Ritwik Singh: Sounds good, Sam. Thanks. Great chatting. Cheers.