AIforIP · Benchmark Collection

Measuring progress in
IP intelligence.

IPIntelligenceArena brings together research-backed benchmarks for evaluating large language models across intellectual property knowledge, practice, and patent examination.

01 / Overview

One field. Three complementary views.

Intellectual property combines legal rules, technical evidence, and professional judgment. The benchmarks in this collection examine distinct but connected capabilities—from applying IP knowledge, to handling diverse practice tasks, to generating examination documents across an iterative patent process.

02 / Benchmarks

Research-backed evaluation,
from knowledge to procedure.

Each benchmark is presented with its original paper, code, data, and project resources.

01
Knowledge & consultation 2024 · Paper

IPEval

A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models

IPEval evaluates how well language models understand and apply intellectual property knowledge in agency and consultation settings. Its bilingual questions span the creation, application, protection, and management of intellectual property.

2,657
Questions
4
Capability dimensions
EN / ZH
Bilingual
02
Comprehensive IP practice ACL 2026

IPBench

Benchmarking Large Language Models on Intellectual Property Knowledge and Practice

IPBench introduces a four-level task taxonomy grounded in Depth of Knowledge. It covers technical and legal tasks across information processing, reasoning, discriminant evaluation, and creative generation in real-world IP practice.

10,374
Data instances
20
Tasks
8
IP mechanisms
03
Patent examination 2026 · Paper

PatRe

A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

PatRe models patent examination as an iterative process of examiner analysis and applicant response. It evaluates Office Action and rebuttal generation under direct, oracle-reference, and retrieval-simulated settings using recent USPTO cases.

480
Real-world cases
1,075
Office Actions
2
Generation tasks

03 / Scope

A structured view of
IP intelligence.

The collection spans complementary evaluation units, task formats, and professional contexts.

Benchmark
Primary focus
Task form
Coverage
IPEvalFoundation
IP knowledge and consultation
Multiple-choice evaluation
English and Chinese · 4 dimensions
IPBenchComprehensive
Knowledge and real-world IP practice
Understanding, reasoning, classification, generation
English and Chinese · 8 IP mechanisms
PatReProcedural
Patent examination lifecycle
Office Action and rebuttal generation
USPTO cases · multi-round histories

04 / Papers

Grounded in public research papers.

Read the original work for task construction, evaluation protocols, experimental settings, findings, and limitations.

2024

IPEval: A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models

Qiyao Wang, Jianguo Huang, Shule Lu, Yuan Lin, Kan Xu, Liang Yang, and Hongfei Lin.

arXiv:2406.12386
2026

Towards IP Intelligence: Benchmarking Large Language Models on Intellectual Property Knowledge and Practice

Qiyao Wang et al. · ACL 2026.

arXiv:2504.15524
2026

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Qiyao Wang, Xinyi Chen, Longze Chen, Hongbo Wang, Hamid Alinejad-Rokny, Yuan Lin, and Min Yang.

arXiv:2605.03571

05 / About

Open research resources for intellectual property intelligence.

IPIntelligenceArena is a curated benchmark collection maintained by AIforIP. It provides a single entry point to research resources that evaluate large language models in intellectual property contexts.

Benchmark descriptions and figures on this site are based on the corresponding papers and public repositories. Please consult the original resources for complete methodological and licensing details.

Visit AIforIP on GitHub