
The legal industry’s relationship with artificial intelligence is seemingly at an inflection point. As a technology first company we firmly believe the use of AI for legal work is not only here to stay, but changing the dynamic of how legal teams work. But, there is still debate and reticence in the legal community about AI use.
Legal has always been a skeptical bunch, and, as a result, the main roadblock to technology adoption has always been trust. For AI that means trusting that the results are accurate, trusting that AI understands the prompt entered and trusting that data is being handled properly.
Some of this mistrust of AI is evidenced by the current cycle of small-scale pilots to determine ROI of AI use, or what some might call “AI theatre”. Despite these pilots, we know legal is moving past the phase of superficial experiments and we know many lawyers and allied professionals are using AI, with or without permission from their organizations.
To date, when legal teams evaluate AI products they must generally rely on benchmarks. These are a useful and needed way to evaluate the accuracy of artificial intelligence for sure. But the question remains, how accurately does a benchmark score really tell you how deeply a model understands legal nuance, it’s “legal intuition”, or more importantly, what type of legal outputs lawyers really prefer in the real world?
Although our core mission remains, to provide legal operations support to corporate legal departments and their law firms, we have also recently embarked on another related mission: To help the legal community and AI companies build artificial intelligence that legal professionals can actually rely on. To do so, we are providing the experience and expertise of our experienced legal professionals to help evaluate and provide feedback for legal AI outputs. As part of this effort, we built Certera.ai.
What is Certera.ai?
Certera.ai is a blind, side-by-side LLM comparison platform built specifically for the legal industry. It brings the concept of blind A/B testing to compare outputs from two anonymized LLM models from legal related prompts generated by lawyers and legal professionals in the wild.
Instead of guessing which commercial or open-source AI model handles specific legal queries best, Certera lets you test them against each other in real time without revealing the models’ identity until after feedback is given on the output.
How It Works
The platform is built on an objective, three-step evaluation process:
Enter a Legal Related Prompt:
Legal professionals enter a real legal related prompt. Before it is sent to any model, users can leverage an integrated data anonymization tool to remove sensitive information.
Review the Blind Outputs:
The query is simultaneously routed to two separate AI models. In anonymous mode, the models used are selected randomly. Or, if the user prefers, two specific models may be selected to compare against each other. Answers populate side-by-side so the user can review and compare the output.
Vote:
Users then apply their professional judgment to select the preferred response based on legal accuracy, structure, and legal reasoning. After the user selects the better response, the vote is recorded and the models are revealed.
The Dynamic Legal AI Leaderboard
Every vote cast by a verified legal professional feeds directly into our public, Legal AI Leaderboard.
Static benchmarks can be artificially gamed if models are trained directly on evaluation data and they rapidly lose utility the moment a new model or iteration is released. By crowdsourcing ground-truth preferences from actual lawyers and allied legal professionals, Certera provides real-time feedback of which models truly lead in legal domain-specific tasks.
Why Percipient Built Certera
For over a decade, we have been using technology to analyze data with humans verifying the output. That is exactly the workflow needed to improve AI output for any discipline.
As a result, we are using our unique position in the legal industry to further augment our technology first mission by helping to ensure the legal tech of the future, which is undoubtedly AI based, is the best it can be. As an added offering we have expanded our services to include providing AI companies with expert evaluation, benchmarking, and RLHF (Reinforcement Learning from Human Feedback) from real-world, experienced legal professionals.
Because our focus is solely on law, we provide an alternative to annotation companies that cater to all industries. Our experience with legal workflows, quality control of legal outputs, and use of experienced legal professionals provides the expertise Frontier AI companies and other enterprises building AI need to ensure that legal outputs are the most accurate possible.
Verifying the true accuracy of legal AI outputs is not best determined by a company’s valuation or marketing hype. It is determined by the precision of the text and analysis on the page.
Ready to see how the top models handle your legal prompts? Head over to Certera.ai, run a few blind tests, and cast your vote to help shape the future of legal technology and artificial intelligence.
The post Judging the Answer, Not the Brand: Why We Built Certera.ai for the Legal Community appeared first on Percipient – Legal Services Powered by Technology.