This repository is dedicated to summarizing papers related to large language models with the field of law
333
95 commits
updated Sep 18, 2026
This repository tracks papers and resources about large language models (LLMs) in the legal domain.
Daily update 2026-09-18 (cs.CL): scanned 104 papers (date=2026-09-18), 2 new papers met the strict "legal task + LLM semantics" inclusion rule.
Category increments today — Applications +2, Legal Reasoning Models +0, Legal Agent +0, Legal Problems +2, Data Resources +2, Law LLMs +0, Evaluation +2.
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) paper
NepKANUN: A RAG-Based Nepali Legal Assistant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records paper
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Privacy Washing: Detecting Internal Contradictions in Privacy Policies paper
Polish ModernBERT: The Long and Short of Polish Language Understanding paper
Annotated Surrogate Retrieval for Polish Statutory Law paper
Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark paper
Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation paper
JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction paper
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences paper
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval paper
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers paper
Modeling Claim Dependency Structure for Patent Litigation Prediction with Graph Attention Networks paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
Generative Gap Filling paper
Benchmarking Patent Drafting from Inventor-Style Disclosures paper
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations paper
ContractScrub: A benchmark for final review of legal contracts paper
When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models paper
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos paper
GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval paper
Redakto - The Incognito Tab for LLMs paper
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Time as Structure: Temporal Dependency Graphs for Verifiable Deadline Computation over Legal Documents paper
How Much Do Legal RAG Systems Still Hallucinate? paper
On Measuring Semantic Preservation in Legal Ontology Learning paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs paper
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? paper
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance paper
Legal Prompt Engineering for Multilingual Legal Judgement Prediction
Can GPT-3 Perform Statutory Reasoning?
Legal Prompting: Teaching a Language Model to Think Like a Lawyer
Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards
ChatGPT Goes to Law School
ChatGPT, Professor of Law
ChatGPT & Generative AI Systems as Quasi-Expert Legal Advice Lawyers - Case Study Considering Potential Appeal Against Conviction of Tom Hayes
‘Words Are Flowing Out Like Endless Rain Into a Paper Cup’: ChatGPT & Law School Assessments
ChatGPT by OpenAI: The End of Litigation Lawyers?
Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans
ChatGPT may Pass the Bar Exam soon, but has a Long Way to Go for the LexGLUE benchmark paper
How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization? paper
Explaining Legal Concepts with Augmented Large Language Models (GPT-4) paper
Garbage in, garbage out: Zero-shot detection of crime using Large Language Models paper
Legal Summarisation through LLMs: The PRODIGIT Project paper
Black-Box Analysis: GPTs Across Time in Legal Textual Entailment Task paper
PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models
Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise paper
Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model Collaboration paper
From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems paper
Boosting legal case retrieval by query content selection with large language models paper
LLMediator: GPT-4 Assisted Online Dispute Resolution paper
Employing Label Models on ChatGPT Answers Improves Legal Text Entailment Performance paper
LLaMandement: Large Language Models for Summarization of French Legislative Proposals paper
Logic Rules as Explanations for Legal Case Retrieval paper Our new paper, welcome to pay attention !!!
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models paper
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models paper
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law paper
Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure paper
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs paper
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model paper
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning paper
GOLDCOIN: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory paper
Enabling Discriminative Reasoning in Large Language Models for Legal Judgment Prediction paper
Large Language Models for Judicial Entity Extraction: A Comparative Study paper
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization paper
LawLLM: Law Large Language Model for the US Legal System paper
Legal syllogism prompting: Teaching large language models for legal judgment prediction paper
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models paper
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools paper
KRAG Framework for Enhancing LLMs in the Legal Domain paper
Legal Evaluations and Challenges of Large Language Models paper
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice paper
Automating Legal Concept Interpretation with LLMs: Retrieval, Generation, and Evaluation paper
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains paper
RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity paper
Courtroom-LLM: A Legal-Inspired Multi-LLM Framework for Resolving Ambiguous Text Classifications paper
LegalViz: Legal Text Visualization by Text To Diagram Generation paper
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning paper
Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning paper
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System paper
LARGE: Legal Retrieval Augmented Generation Evaluation Tool paper
Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification paper
The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting paper
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech paper
CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks paper
LLM-based Embedders for Prior Case Retrieval paper
Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models paper
Using Large Language Models for Legal Decision-Making in Austrian Value-Added Tax Law: An Experimental Study paper
ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation paper
LLMs in Interpreting Legal Documents paper
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment paper
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics paper
Jurisdiction as Structural Barrier: How Privacy Policy Organization May Reduce Visibility of Substantive Disclosures paper
Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry paper
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model paper
Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic Languages paper
AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments paper
LLM-Assisted Causal Structure Disambiguation and Factor Extraction for Legal Judgment Prediction paper
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules paper
ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance paper
Eskwai for Students: Generative AI Assistant for Legal Education in Ghana paper
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization paper
Exploring Lightweight Large Language Models for Court View Generation paper
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free paper
Chunking German Legal Code paper
Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation paper
Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval paper
Retrieval-Augmented Detection of Potentially Abusive Clauses in Chilean Terms of Service paper
Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis paper
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization paper
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper
Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper
EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper
DAR: Deontic Reasoning with Agentic Harnesses paper
GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation paper
Civil Court Simulation with Large Language Models paper
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper
FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs paper
Reinforcement learning to improve large language model-based automated code compliance systems paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs paper
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation paper
Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing paper
When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation paper
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation paper
Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning paper
AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System paper
Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps paper
From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE paper
Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management paper
Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh paper
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
ContractScrub: A benchmark for final review of legal contracts paper
Redakto - The Incognito Tab for LLMs paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks paper
How Much Do Legal RAG Systems Still Hallucinate? paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Is this Citation on Point? paper
SoK: From Generation to Consumption of Privacy Documents in Software Systems paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
LexKairos: Benchmarking Legal Temporal Capabilities in LLMs paper
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models
Persistent Anti-Muslim Bias in Large Language Models
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination
The GPTJudge: Justice in a Generative AI World paper
Is the U.S. Legal System Ready for AI's Challenges to Human Values? paper
Questioning Biases in Case Judgment Summaries: Legal Datasets or Large Language Models? paper
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
A Legal Framework for Natural Language Processing Model Training in Portugal paper
LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain paper
Bias in Large Language Models: Origin, Evaluation, and Mitigation paper
An Information Theoretic Approach to Operationalize Right to Data Protection paper
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench paper
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance paper
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law paper
Large Language Models as Search Engines: Societal Challenges paper
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models paper
TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC paper
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model paper
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Who Checks the Citations? Benchmarking Legal Hallucination Detection paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
Position: The Term "Machine Unlearning" Is Overused in LLMs paper
Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages paper
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper
Agentic Evaluation of Copyright Law Compliance paper
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant paper
Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) paper
NepKANUN: A RAG-Based Nepali Legal Assistant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records paper
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation paper
Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Polish ModernBERT: The Long and Short of Polish Language Understanding paper
Annotated Surrogate Retrieval for Polish Statutory Law paper
Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark paper
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences paper
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval paper
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers paper
Modeling Claim Dependency Structure for Patent Litigation Prediction with Graph Attention Networks paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
Generative Gap Filling paper
Benchmarking Patent Drafting from Inventor-Style Disclosures paper
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations paper
ContractScrub: A benchmark for final review of legal contracts paper
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos paper
GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval paper
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks paper
Time as Structure: Temporal Dependency Graphs for Verifiable Deadline Computation over Legal Documents paper
How Much Do Legal RAG Systems Still Hallucinate? paper
On Measuring Semantic Preservation in Legal Ontology Learning paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Is this Citation on Point? paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
LexKairos: Benchmarking Legal Temporal Capabilities in LLMs paper
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs paper
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts paper
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? paper
LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports paper
Measuring Massive Multitask Chinese Understanding paper
LawBench: Benchmarking Legal Knowledge of Large Language Models github
Large Language Models are legal but they are not: Making the case for a powerful LegalLLM paper
Better Call GPT, Comparing Large Language Models Against Lawyers paper
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics paper
Evaluation Ethics of LLMs in Legal Domain paper
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination paper
LawBench: Benchmarking Legal Knowledge of Large Language Models paper
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models paper
LegalAgentBench: Evaluating LLM Agents in Legal Domain paper
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments paper
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text paper
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning paper
LexTime: A Benchmark for Temporal Ordering of Legal Events paper
LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements paper
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering paper
The Massive Legal Embedding Benchmark (MLEB) paper
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence paper
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents paper
Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Validation paper
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text paper
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization paper
Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study paper
Reasoners or Translators? Contamination-aware Evaluation and Neuro-Symbolic Robustness in Tax Law paper
Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification paper
LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation paper
GradeLegal: Automated Grading for German Legal Cases paper
Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering paper
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays paper
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment paper
The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment paper
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law paper
Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions paper
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning paper
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification paper
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text paper
RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora paper
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper
Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper
EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper
HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule paper
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection paper
DeXposure-Claw: An Agentic System for DeFi Risk Supervision paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Who Checks the Citations? Benchmarking Legal Hallucination Detection paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper
Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR paper
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper
The survey paper is shown in paper
Please cite the following papers as references if this repository helps you (*ゝω・)ノThanks!
@article{sun2023short,
title={A short survey of viewing large language models in legal aspect},
author={Sun, Zhongxiang},
journal={arXiv preprint arXiv:2303.09136v2},
year={2023}
}
95 commits
This repository is dedicated to summarizing papers related to large language models with the field of law
333
95 commits
updated Sep 18, 2026
This repository tracks papers and resources about large language models (LLMs) in the legal domain.
Daily update 2026-09-18 (cs.CL): scanned 104 papers (date=2026-09-18), 2 new papers met the strict "legal task + LLM semantics" inclusion rule.
Category increments today — Applications +2, Legal Reasoning Models +0, Legal Agent +0, Legal Problems +2, Data Resources +2, Law LLMs +0, Evaluation +2.
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) paper
NepKANUN: A RAG-Based Nepali Legal Assistant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records paper
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Privacy Washing: Detecting Internal Contradictions in Privacy Policies paper
Polish ModernBERT: The Long and Short of Polish Language Understanding paper
Annotated Surrogate Retrieval for Polish Statutory Law paper
Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark paper
Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation paper
JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction paper
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences paper
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval paper
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers paper
Modeling Claim Dependency Structure for Patent Litigation Prediction with Graph Attention Networks paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
Generative Gap Filling paper
Benchmarking Patent Drafting from Inventor-Style Disclosures paper
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations paper
ContractScrub: A benchmark for final review of legal contracts paper
When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models paper
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos paper
GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval paper
Redakto - The Incognito Tab for LLMs paper
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Time as Structure: Temporal Dependency Graphs for Verifiable Deadline Computation over Legal Documents paper
How Much Do Legal RAG Systems Still Hallucinate? paper
On Measuring Semantic Preservation in Legal Ontology Learning paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs paper
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? paper
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance paper
Legal Prompt Engineering for Multilingual Legal Judgement Prediction
Can GPT-3 Perform Statutory Reasoning?
Legal Prompting: Teaching a Language Model to Think Like a Lawyer
Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards
ChatGPT Goes to Law School
ChatGPT, Professor of Law
ChatGPT & Generative AI Systems as Quasi-Expert Legal Advice Lawyers - Case Study Considering Potential Appeal Against Conviction of Tom Hayes
‘Words Are Flowing Out Like Endless Rain Into a Paper Cup’: ChatGPT & Law School Assessments
ChatGPT by OpenAI: The End of Litigation Lawyers?
Law Informs Code: A Legal Informatics Approach to Aligning Artificial Intelligence with Humans
ChatGPT may Pass the Bar Exam soon, but has a Long Way to Go for the LexGLUE benchmark paper
How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization? paper
Explaining Legal Concepts with Augmented Large Language Models (GPT-4) paper
Garbage in, garbage out: Zero-shot detection of crime using Large Language Models paper
Legal Summarisation through LLMs: The PRODIGIT Project paper
Black-Box Analysis: GPTs Across Time in Legal Textual Entailment Task paper
PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models
Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise paper
Precedent-Enhanced Legal Judgment Prediction with LLM and Domain-Model Collaboration paper
From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems paper
Boosting legal case retrieval by query content selection with large language models paper
LLMediator: GPT-4 Assisted Online Dispute Resolution paper
Employing Label Models on ChatGPT Answers Improves Legal Text Entailment Performance paper
LLaMandement: Large Language Models for Summarization of French Legislative Proposals paper
Logic Rules as Explanations for Legal Case Retrieval paper Our new paper, welcome to pay attention !!!
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models paper
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models paper
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law paper
Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure paper
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs paper
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model paper
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning paper
GOLDCOIN: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory paper
Enabling Discriminative Reasoning in Large Language Models for Legal Judgment Prediction paper
Large Language Models for Judicial Entity Extraction: A Comparative Study paper
Applicability of Large Language Models and Generative Models for Legal Case Judgement Summarization paper
LawLLM: Law Large Language Model for the US Legal System paper
Legal syllogism prompting: Teaching large language models for legal judgment prediction paper
Optimizing Numerical Estimation and Operational Efficiency in the Legal Domain through Large Language Models paper
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools paper
KRAG Framework for Enhancing LLMs in the Legal Domain paper
Legal Evaluations and Challenges of Large Language Models paper
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice paper
Automating Legal Concept Interpretation with LLMs: Retrieval, Generation, and Evaluation paper
Domaino1s: Guiding LLM Reasoning for Explainable Answers in High-Stakes Domains paper
RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity paper
Courtroom-LLM: A Legal-Inspired Multi-LLM Framework for Resolving Ambiguous Text Classifications paper
LegalViz: Legal Text Visualization by Text To Diagram Generation paper
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning paper
Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning paper
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System paper
LARGE: Legal Retrieval Augmented Generation Evaluation Tool paper
Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification paper
The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting paper
Conditioning Large Language Models on Legal Systems? Detecting Punishable Hate Speech paper
CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks paper
LLM-based Embedders for Prior Case Retrieval paper
Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models paper
Using Large Language Models for Legal Decision-Making in Austrian Value-Added Tax Law: An Experimental Study paper
ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation paper
LLMs in Interpreting Legal Documents paper
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment paper
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics paper
Jurisdiction as Structural Barrier: How Privacy Policy Organization May Reduce Visibility of Substantive Disclosures paper
Enforcing Monotonic Progress in Legal Cross-Examination: Preventing Long-Horizon Stagnation in LLM-Based Inquiry paper
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model paper
Synthesizing the Virtual Advocate: A Multi-Persona Speech Generation Framework for Diverse Linguistic Jurisdictions in Indic Languages paper
AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments paper
LLM-Assisted Causal Structure Disambiguation and Factor Extraction for Legal Judgment Prediction paper
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules paper
ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance paper
Eskwai for Students: Generative AI Assistant for Legal Education in Ghana paper
Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text Summarization paper
Exploring Lightweight Large Language Models for Court View Generation paper
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free paper
Chunking German Legal Code paper
Generating Legal Commentaries from Case Databases via Retrieval, Clustering, and Generation paper
Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval paper
Retrieval-Augmented Detection of Potentially Abusive Clauses in Chilean Terms of Service paper
Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis paper
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization paper
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper
Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper
EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper
DAR: Deontic Reasoning with Agentic Harnesses paper
GARL: Game-Theoretic Reinforcement Learning for Multi-Agent Strategic Prioritisation paper
Civil Court Simulation with Large Language Models paper
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper
FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs paper
Reinforcement learning to improve large language model-based automated code compliance systems paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs paper
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation paper
Cross-Architecture LLM Ensembles, Feature-Based Reranking and Retrieval-Augmented Prompting for Legal Information Processing paper
When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation paper
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation paper
Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning paper
AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System paper
Enabling Multilingual Privacy Policy Audits: Large-Scale Analysis of Spanish Mobile Apps paper
From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE paper
Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management paper
Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh paper
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
ContractScrub: A benchmark for final review of legal contracts paper
Redakto - The Incognito Tab for LLMs paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks paper
How Much Do Legal RAG Systems Still Hallucinate? paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Is this Citation on Point? paper
SoK: From Generation to Consumption of Privacy Documents in Software Systems paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
LexKairos: Benchmarking Legal Temporal Capabilities in LLMs paper
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models
Persistent Anti-Muslim Bias in Large Language Models
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
The Dark Side of ChatGPT: Legal and Ethical Challenges from Stochastic Parrots and Hallucination
The GPTJudge: Justice in a Generative AI World paper
Is the U.S. Legal System Ready for AI's Challenges to Human Values? paper
Questioning Biases in Case Judgment Summaries: Legal Datasets or Large Language Models? paper
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
A Legal Framework for Natural Language Processing Model Training in Portugal paper
LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain paper
Bias in Large Language Models: Origin, Evaluation, and Mitigation paper
An Information Theoretic Approach to Operationalize Right to Data Protection paper
Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench paper
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance paper
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law paper
Large Language Models as Search Engines: Societal Challenges paper
From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models paper
TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC paper
To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model paper
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Who Checks the Citations? Benchmarking Legal Hallucination Detection paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
Position: The Term "Machine Unlearning" Is Overused in LLMs paper
Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages paper
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper
Agentic Evaluation of Copyright Law Compliance paper
Benchmarking LLM Compliance with China AI Generated Content Regulations paper
Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence paper
How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards paper
Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment paper
Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant paper
Nepali Legal Expertise through Generative and Extractive Pre-trained Transformers (NepLEGiT) paper
NepKANUN: A RAG-Based Nepali Legal Assistant paper
Few-Shot Degradation Is Not What It Seems: Behavioral Evidence, Representation Analysis, and a Random-Text Control Across 12 Models, 2 Tasks, and 2 Architectures paper
CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering paper
Building Legal Reward Models for Grounding and Abstention paper
Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging paper
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models paper
ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation paper
BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation paper
GANDR: Claim Auditing for Verifiable Legal Answer Generation paper
When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors paper
MUCnoHARM@GermEval Shared Task 2026: Retrieval-based In-Context Learning for Defamatory Offences, and Where It Falls Short paper
Reading a Legal Question Word by Word: Embedding Trajectories of 2,144 Vietnamese Legal Headlines paper
FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect paper
Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems paper
LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics paper
KhatianDoc: A Human-Verified Benchmark Diagnosing Multimodal LLM Failure on Bengali Legal Land Records paper
LexIssue: Benchmarking Legal Issue Identification in Chinese Civil Litigation paper
Rent-a-RAG: Embedding-Space Watermarks for Auditing Third-Party RAG paper
OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction paper
Polish ModernBERT: The Long and Short of Polish Language Understanding paper
Annotated Surrogate Retrieval for Polish Statutory Law paper
Do Small Models Use the Law You Give Them? Measuring Context Use on a Bilingual Bangladesh Legal Benchmark paper
Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning paper
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict paper
RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences paper
GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval paper
Figurative Justice: Detecting metaphors in Hindi judgements with qualitative assessment and transformers paper
Modeling Claim Dependency Structure for Patent Litigation Prediction with Graph Attention Networks paper
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? paper
Generative Gap Filling paper
Benchmarking Patent Drafting from Inventor-Style Disclosures paper
Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations paper
ContractScrub: A benchmark for final review of legal contracts paper
ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos paper
GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval paper
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method paper
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases paper
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion paper
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks paper
Time as Structure: Temporal Dependency Graphs for Verifiable Deadline Computation over Legal Documents paper
How Much Do Legal RAG Systems Still Hallucinate? paper
On Measuring Semantic Preservation in Legal Ontology Learning paper
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance paper
Is this Citation on Point? paper
Evaluating Rational Contracting in Natural Language paper
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law paper
LexKairos: Benchmarking Legal Temporal Capabilities in LLMs paper
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs paper
ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts paper
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks paper
CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models paper
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art paper
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? paper
LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports paper
Measuring Massive Multitask Chinese Understanding paper
LawBench: Benchmarking Legal Knowledge of Large Language Models github
Large Language Models are legal but they are not: Making the case for a powerful LegalLLM paper
Better Call GPT, Comparing Large Language Models Against Lawyers paper
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics paper
Evaluation Ethics of LLMs in Legal Domain paper
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination paper
LawBench: Benchmarking Legal Knowledge of Large Language Models paper
LexEval: A Comprehensive Chinese Legal Benchmark for Evaluating Large Language Models paper
LegalAgentBench: Evaluating LLM Agents in Legal Domain paper
Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments paper
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text paper
When Fairness Isn't Statistical: The Limits of Machine Learning in Evaluating Legal Reasoning paper
LexTime: A Benchmark for Temporal Ordering of Legal Events paper
LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements paper
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering paper
The Massive Legal Embedding Benchmark (MLEB) paper
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence paper
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents paper
Adaptive Cost-Efficient Evaluation for Reliable Patent Claim Validation paper
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text paper
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization paper
Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study paper
Reasoners or Translators? Contamination-aware Evaluation and Neuro-Symbolic Robustness in Tax Law paper
Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification paper
LP-Eval: Rubric and Dataset for Measuring the Quality of Legal Proposition Generation paper
GradeLegal: Automated Grading for German Legal Cases paper
Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering paper
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays paper
JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment paper
The Cases LJP Never Sees: Prosecution Decision Prediction for More Complete Criminal Liability Assessment paper
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law paper
Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions paper
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning paper
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification paper
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text paper
RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora paper
ImmigrationQA: A Source-Grounded Dataset and Small-Model Adaptation for U.S. Immigration Law paper
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law paper
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs paper
Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs paper
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral paper
Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA paper
EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction paper
HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule paper
LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks paper
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing paper
Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering paper
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection paper
DeXposure-Claw: An Agentic System for DeFi Risk Supervision paper
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts paper
Who Checks the Citations? Benchmarking Legal Hallucination Detection paper
Peeking Inside LLMs: Leveraging Internal Artifacts of LLMs for Enhancing Reliability in Legal Classification paper
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law paper
Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR paper
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI paper
The survey paper is shown in paper
Please cite the following papers as references if this repository helps you (*ゝω・)ノThanks!
@article{sun2023short,
title={A short survey of viewing large language models in legal aspect},
author={Sun, Zhongxiang},
journal={arXiv preprint arXiv:2303.09136v2},
year={2023}
}
95 commits