Legal research relies on accurate information. A missing detail, unsupported statement, or incorrect answer can lead to flawed legal reasoning.
As part of the Global AI Assurance Sandbox, led by IMDA and the AI Verify Foundation, Resaro conducted independent testing of LawNet AI, a Q&A system built by the Singapore Academy of Law (SAL) to support legal research on its LawNet platform.
Using Resaro's Approved Intelligence Platform (AIP) and its RAG test plan, testing centred on three prioritised risks:
-
Correctness of answers, including faithfulness and lack of completeness. LawNet AI could generate incorrect, misleading, or incomplete legal answers that omit important legal qualifications or fabricate unsupported claims, which could negatively affect lawyers’ research quality, reduce trust in the platform, and potentially impact legal decision-making.
-
Data leakage, specifically system prompt leakage. SAL wanted to ensure that the system would not inadvertently disclose confidential, sensitive, or internal information such as system prompts.
-
Reliability (consistency of responses for identical prompts) and robustness to perturbations. SAL needed the system to provide stable and consistent legal answers to the same queries, as well as to meaning-preserving perturbations of queries (typos, synonyms, filler words, paraphrasing, and tones).
Key Insights
-
General-purpose synthetic data and testing pipelines are not built for specialised domains. Legal concepts can be phrased many different ways, and older case law can be superseded by newer decisions. A testing approach that doesn't account for this will miss failures that a legal expert would catch immediately.
-
Precision of context is critical to reliable evaluation. Small changes in how a legal question is framed can shift what the correct answer actually is. Datasets, whether human-annotated or synthetic, need enough context to anchor a single clear reference answer, otherwise the evaluation results themselves become unreliable.
-
Correctness alone is not enough to evaluate a legal Q&A system. Faithfulness, citation grounding, and completeness all needed to be assessed together to get a reliable picture of whether LawNet AI's answers could be trusted for real legal research.
Access the full case study here:
https://file.go.gov.sg/sandbox-salxresaro.pdf
More about Resaro, and our role in the Global AI Assurance Sandbox
Resaro is an AI assurance company that provides the testing and validation stack for mission-critical AI deployment, enabling confident use of AI beyond the lab. Co-headquartered in Singapore and Munich, Resaro works with defence, government, and critical infrastructure organisations to move AI systems from pilots into real-world operations.
Resaro’s Approved Intelligence Platform (AIP) gives engineers and operators the control and evidence needed to move systems from idea to mission-ready. The platform is built around scenario-specific synthetic data pipelines mapped to mission-defined test plans, recreating operational conditions, edge cases, and failure scenarios that determine safe and effective deployment.
This enables validation across the AI lifecycle, producing structured, auditable evidence of performance, limitations, and failure behaviour. The evidence is consolidated into an AI Solutions Quality Index, giving leadership teams a shared, decision-ready view of deployment readiness for confident, controlled adoption.
In the Global AI Assurance Sandbox, our role was to assess the deployment risks of each GenAI application through architectural review and stakeholder collaboration, define the use-case-specific test plans that mattered for each system, and run the scenario-specific testing needed to produce evidence that reflects real operational conditions, not lab conditions. This spanned scenario-based synthetic data generation, automated LLM-as-a-judge evaluation, and human expert validation, consolidated into an AI Solutions Quality Index for each application.