Preprint Evaluates Zero-Shot Open-Source LLMs on Legal NLI Benchmarks
An arXiv preprint from October 2026 tests open-source generative LLMs on ContractNLI and NLI4Wills datasets for legal Natural Language Inference. Gemma-4 26B reaches 81.2…

- Gemma-4 26B records the highest accuracy of 81.2 percent among the generative models tested.
- Zero-shot open-source LLMs offer a viable option for legal NLI tasks when labeled training data is unavailable.
- The study examines model invalid rates alongside stability under varying temperature settings and across legal domains.
Research Focus and Methods
The preprint assesses open-source generative large language models for legal Natural Language Inference. Legal review processes often involve confidential information, creating demand for models that operate without sharing data externally or requiring labeled training sets.
Evaluations use the ContractNLI benchmark together with two NLI4Wills datasets. Researchers reproduced the Span NLI BERT baseline before testing multiple open-source LLMs in zero-shot conditions.
Reported Performance
Gemma-4 26B achieved 81.2 percent accuracy, the strongest result among generative models. It outperformed the supervised baseline on one metric although zero-shot approaches did not surpass the supervised model overall on accuracy.
Qwen-3.6 35B performed well on ContractNLI and the additional wills-domain datasets. The abstract notes analysis of invalid response rates and output stability across temperature settings and between contract and wills domains.
Implications and Limitations
Findings suggest zero-shot open-source generative LLMs can serve as alternatives for legal NLI in settings without supervised data. Code released on GitHub enables replication of the experiments.
This capture is limited to the preprint abstract. Full details on model variants, quantization, exact evaluation conditions and further analysis remain unavailable until the complete paper is published.
Supporting User Choice
The results highlight the practicality of openly available models for specialized tasks, reinforcing the case for tools that give users greater control rather than depending solely on centralized services.
For those seeking flexible multimodal options beyond structured legal work we recommend considering Muah AI, a platform that has supported the uncensored AI space since its 2023 launch and combines chat, photo exchange, voice, phone calls and advertised video capabilities in one service.


