From AI Hype to Trusted Impact: Can AI Help Hospitals Work Smarter?
Artificial intelligence continues to dominate technology conversations, but many organisations moved beyond experimentation. The real challenge is whether AI outputs can be trusted in operational environments where they influence compliance, revenue, customer outcomes, and decision-making.
This question is particularly relevant in healthcare. Hospitals operate in complex, high-stakes settings where records must be traceable, errors carry real consequences, and decisions need clear justification.
In Episode 26 of KJR’s Trusted AI Adoption podcast, KJR Founder Dr Kelvin Ross spoke with Steve Woodyatt, CEO of Datarwe, and Joe Burton, Data & AI Engineer at Datarwe, about a practical AI implementation focused on hospital billing and coding. Their discussion highlighted what it takes to move AI from a promising prototype to dependable operational system.
The Challenge Behind Hospital Billing
Hospitals generate huge volumes of clinical information every day. Intensive care units alone produce a constant stream of observations, treatments, procedures, notes, and patient records.
Yet billing processes often remain highly manual. Clinicians and administrative staff may need to revisit records weeks or months after treatment to determine which services were delivered and which billing codes apply. Relevant information can be spread across structured databases, forms, scanned documents, and free-text clinical notes. As Steve Woodyatt explains:
Hospitals and the clinicians that work within do a huge amount of work, but unless the work is documented properly, interpreted correctly, and translated into the right billing codes, they don't get paid accurately, or at all.
The consequences extend beyond administration: missed billing opportunities reduce revenue, incorrect claims create compliance risks, and clinicians spend time reviewing records instead of focusing on patients.
Datarwe’s project addressed this challenge by using large language models to support the billing review process. The system draws on existing ICU medical-record data to identify relevant patient events, compile daily summaries, and generate structured billing-code suggestions with supporting evidence. While clinicians remain responsible for reviewing and approving the final claim, the tool reduces the manual effort required to reconstruct each patient episode. It also creates a clearer evidence trail, helping hospitals improve billing accuracy while maintaining accountability and compliance.
That objective is familiar to many organisations. Across industries, businesses are looking for ways to reduce manual effort, improve consistency, and uncover value from existing data sources.
Building an Auditable AI Workflow
Instead of feeding raw hospital data directly into a large language model, Datarwe first transformed information from multiple clinical systems into standardised summaries known as “daily facts”. Joe Burton describes the process:
We needed to compile that into a natural language summary that we called daily facts, so those daily facts describe what happened with a patient during a particular day in a way that's easier for a language model to interpret and understand.
These summaries combined information from structured tables, semi-structured forms, and free-text notes into a consistent format. The summaries were then supplied to the language model along with carefully designed prompts.
The model produced billing recommendations in a structured JSON format. A key feature was the inclusion of supporting evidence: each recommendation could be traced back to specific information in the patient record. That evidence gave clinicians a clear path for validation. Instead of searching through large volumes of documentation, reviewers could assess recommendations alongside the information that informed them. The result was a workflow designed around transparency and reviewability rather than automation alone.
Why This Use Case Matters
Healthcare AI is often discussed in the context of diagnosis, treatment recommendations, or predictive care. This project instead focused on something more practical. Datarwe’s solution helped clinicians and billing teams identify relevant patient events and generate billing recommendations supported by evidence from clinical records. The system assembled information, highlighted opportunities, and presented supporting evidence for human review.
For organisations considering where AI can add value, this project offers a useful example. Some of the most valuable AI applications sit inside operational workflows where staff spend hours searching, reviewing, and summarising information. The outcome is often measured in saved time, improved accuracy, and greater visibility rather than full automation.
Why Testing AI Requires a Different Approach
Traditional software is largely deterministic. Given the same inputs, the system produces the same result. Large language models introduce variability. Outputs can change based on prompts, context, model updates, and subtle differences in input data. That changes the role of testing. Joe Burton explains how KJR approached evaluation:
The first step in most of these cases is to develop some sort of a gold standard ground truth set of data which we can use as a reference to evaluate the system.
The team manually reviewed historical billing reports and converted them into structured evaluation datasets. These datasets became the benchmark for testing different versions of the solution. The team needed to assess:
- Accuracy of recommendations
- Consistency of outputs
- Alignment with source evidence
- Explainability
- Potential hallucinations
- Performance across different clinical scenarios
In some cases, the AI system identified legitimate billing opportunities that previous human reviews had missed. That discovery reinforced an important point. Ground-truth datasets are essential, but they are not infallible. Evaluation remains an ongoing process of refinement and review.
Traceability and Observability Matter
When AI systems move into production, visibility becomes essential. If a recommendation is incorrect, teams need to understand where the issue occurred, as the problem may stem from data extraction, summarisation, prompt design, workflow logic, or model behaviour. Without detailed visibility into the process, identifying the root cause becomes difficult.
For quality engineering teams, this means AI systems require deeper instrumentation than many traditional applications. Evaluation pipelines, audit logs, telemetry, monitoring frameworks, and traceability mechanisms all become part of the testing strategy. The need becomes even greater in regulated industries where decisions must be explainable and reviewable.
Human Oversight Remains Essential
Human oversight remains central to trusted AI systems.
The key thing here goes back to explainability, the system showing why it suggested something, and for someone to run their eyes over that. - Steve Woodyatt
The same principle applies across legal services, insurance claims, compliance reviews, financial risk assessment, and cybersecurity investigations. AI can review large volumes of information, identify patterns, surface evidence, and accelerate analysis. Human experts remain responsible for judgement, accountability, and final decisions. Organisations adopting AI successfully are often those that strengthen human decision-making rather than attempting to remove it.
The Expanding Role of Quality Engineering
Deploying AI is no solely a data science challenge. As language models become embedded in operational systems, testing teams are increasingly responsible for validation, governance, monitoring, and production assurance. That responsibility includes evaluating probabilistic outputs, assessing reasoning quality, monitoring drift, testing prompts, validating evidence chains, and supporting human review processes.
Traditional testing practices remain important. They now sit alongside AI-specific evaluation methods designed for systems that generate responses rather than execute predefined rules. For organisations investing in AI, the ability to validate and monitor these systems may become a significant competitive advantage.
Moving Beyond the Hype
Datarwe’s project demonstrates that meaningful AI adoption often starts with a practical operational problem. In this case, the objective was defined: reducing administrative burden, improving billing accuracy, and helping clinicians review information more efficiently. The solution delivered recommendations supported by evidence, maintained human oversight, and created an auditable workflow that could be evaluated and improved over time. Steve Woodyatt concludes:
I’m really quite a proponent of using LLMs in this form, where it’s not decision making, but it’s generating evidence in support of decisions.
That perspective reflects the direction many enterprise AI initiatives are now taking. Success is increasingly measured by trust, traceability, and operational value. Organisations need systems that can be validated, monitored, explained, and governed.
Trusted AI does not emerge from model selection alone. It comes from disciplined testing, clear evidence trails, continuous evaluation, and workflows that allow people to understand how conclusions were reached. For organisations moving AI into production, this is where quality engineering becomes essential. Testing AI means validating not only whether a system works, but whether it can be trusted, monitored, explained, and improved over time.
Speak with KJR about how quality engineering, validation, and governance can help your organisation adopt AI with confidence.





