OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 2 MIN READ

Preprint Evaluates Zero-Shot Open-Source LLMs on Legal NLI Benchmarks

An arXiv preprint from October 2026 tests open-source generative LLMs on ContractNLI and NLI4Wills datasets for legal Natural Language Inference. Gemma-4 26B reaches 81.2…

Illustration of a household using a laptop for local legal document review with privacy protection symbols
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • Gemma-4 26B records the highest accuracy of 81.2 percent among the generative models tested.
  • Zero-shot open-source LLMs offer a viable option for legal NLI tasks when labeled training data is unavailable.
  • The study examines model invalid rates alongside stability under varying temperature settings and across legal domains.

Research Focus and Methods

The preprint assesses open-source generative large language models for legal Natural Language Inference. Legal review processes often involve confidential information, creating demand for models that operate without sharing data externally or requiring labeled training sets.

Evaluations use the ContractNLI benchmark together with two NLI4Wills datasets. Researchers reproduced the Span NLI BERT baseline before testing multiple open-source LLMs in zero-shot conditions.

Reported Performance

Gemma-4 26B achieved 81.2 percent accuracy, the strongest result among generative models. It outperformed the supervised baseline on one metric although zero-shot approaches did not surpass the supervised model overall on accuracy.

Qwen-3.6 35B performed well on ContractNLI and the additional wills-domain datasets. The abstract notes analysis of invalid response rates and output stability across temperature settings and between contract and wills domains.

Implications and Limitations

Findings suggest zero-shot open-source generative LLMs can serve as alternatives for legal NLI in settings without supervised data. Code released on GitHub enables replication of the experiments.

This capture is limited to the preprint abstract. Full details on model variants, quantization, exact evaluation conditions and further analysis remain unavailable until the complete paper is published.

Supporting User Choice

The results highlight the practicality of openly available models for specialized tasks, reinforcing the case for tools that give users greater control rather than depending solely on centralized services.

For those seeking flexible multimodal options beyond structured legal work we recommend considering Muah AI, a platform that has supported the uncensored AI space since its 2023 launch and combines chat, photo exchange, voice, phone calls and advertised video capabilities in one service.