articledataartificial intelligencelegal practitionerlegal ai

Why data is the real foundation

The internal debate within our team has oscillated between the wish to convey to users the work we’ve been doing with data since 2019 and the alternate view: do people really care about data-hygiene and standards? Are we not conditioned at this stage to just expect the data to be there for any LLM model to use?

In this first instalment, we outline the data imperative:

The AI Assumption

Every conversation about legal technology now starts with AI — new entrants, new products, new pitches, almost all framed around what a model can do. The data the model is reasoning over is treated as a solved problem.

It isn’t. Vizlegal’s bet, since 2019, has been the opposite: mind the data, mine the data.

Search, Timelines, MyCalendar, Firm and Judge Profiles — and the forthcoming Query Builder and Smart Assistant — only work because of the six years we’ve spent on the structured spine beneath them.

A Founding Observation

The realisation in 2019 was simple: publicly held legal data was catalogued, but it wasn’t structured. A judgment, a High Court listing, a Labour Court determination could be located — but its dates, parties, legislative provisions, outcomes and adjudicators were not uniformly presented.

The cost of that, and the risk of omission or misinterpretation, fell on the practitioner. It also, more quietly, compromised access to justice itself. Anyone whose work depended on tracing a matter through the system was working against the data rather than with it.

That was the gap we set out to close. The application layer came afterwards.

The Application-Layer Distinction

This is where Vizlegal’s path diverges from the cohort of products the market now collectively calls “legal AI” or application-layer products.  In the context of their research capability they place their model in front of structured data that is typically rented through partnerships and tie-ins; the longevity and breadth of which the end user should be vigilant about.

The local & specialist appreciation for how sources actually behave is of immediate relevance to a legal profession leveraged on accurate information. The model sits on top of someone else’s corpus, and any peculiarity in that collection - a missing identifier, a retired source, an inconsistent reference convention - is invisible to the layer above.

💡

Would you like to be included in the testing of future Vizlegal features? We’re interested in building a cohort of innovation-focussed practitioners looking to responsibly integrate data and AI based tools to support their practice.

Get in touch here and we’ll include you in our next round of feature roll out.

Data is never straightforward

There is a tempting fiction that public legal and regulatory data sits on government websites in a form a system can simply ingest.

Public legal data magnifies the point. Conventions shift across years and decades.. Sources change formats without warning. Authorities retire pages and archive whole sections of their websites — and what they retire, generalist AI models cannot necessarily see.

A small example: Workplace Relations data in its rawest sense never indicates where a decision has been appealed to the Labour Court, and neither is it searchable as a single collection by keyword, party or phrase. Our expertise and experience delves deeper, to link WRC cases to a Labour Court Adjudication, on all the grounds of appeal.  We do the same for our 590k High Court Records, corresponding judgments and appeals.

A small omission can have large downstream consequences for anyone tracing a matter through the system.

That is editorial work that we at Vizlegal lean into.  As our CTO José Alberto Suárez López attests:

“People assume the hard part of legal AI is the AI. It isn't. The hard part is sitting with a single source for weeks until you actually understand how it behaves — and then doing that thirty-six more times.
There is no shortcut to understanding a legal source. You read it, you talk to the people who produce it, you map its quirks, and then you build around what you've learned. That work is what gives our users something they can trust”.

AI Results: A function of the data it is reasoning over

The Irish judiciary’s June 2024 guidance to judges put the matter plainly, in that the outcome of generative AI models ‘cannot be restricted to provide answers solely from authoritative databases, or databases with any significant training data on Irish law.

The UK Judicial Office reached the same conclusion in its October 2025 guidance. The reliability of any AI surface - judicial, professional or commercial - is a function of the data it is reasoning over.

If the corpus is structured, curated and traceable, it can be relied on. The model is the easy part.

Vizlegal’s daily ingestion and arrangement of 36 sources, 5m+ documents to-date,  across three legal jurisdictions, ranging from judicial to quasi-judicial to administrative data; are not simply ‘added’, they are weeded and graded, tested and cleaned so that users have confidence in the features that they underpin.

What this looks like in practice

One of the features we are soon to release - working title Smart Assistant - is an example what our technical team is building on the shoulders of the data they have so carefully maintained.

When a user asks our Smart Assistant a question “What is the most precedential case in Ireland on the definition of public authority under the Aarhus Convention and why?”; the system does not improvise. It generates structured queries against named fields, with explicit synonyms, related concepts, and filters bound to the court hierarchy. The user can inspect every query that was run and understand the path taken.

That transparency is not cosmetic. It is the only reason the answer can be trusted. And it is only possible because, beneath the AI surface, every document in our corpus has been ingested, classified, normalised, validated and indexed  by our technical team over years.

The series ahead

Forthcoming instalments will go deeper on our data-cleansing approach and the role of practitioners in it; preview the Query Builder, our most sophisticated research tool yet, due in beta in the coming months; and return to the Smart Assistant and the trajectory beyond it — AI built on structured data.

The order matters. The data has to come first. It always did.

💡
Building AI literacy and a culture of innovation? Invite us to present a CPD to your colleagues.


Book a live demo

Book a live demo