AIforIP · Benchmark Collection

Measuring progress in
IP Intelligence.

IPIntelligenceArena brings together research-backed benchmarks for evaluating large language models across intellectual property knowledge, practice, and patent examination.

01 / Overview

From benchmark questions to professional workflows.

IPIntelligenceArena organizes the team’s public benchmarks along a progression of IP capabilities. IPEval tests bilingual knowledge and consultation, IPBench expands evaluation to legal and technical practice tasks, and PatRe moves into multi-round patent examination, where evidence, procedural context, and document generation must work together.

Together, they show how evaluation is moving from knowing IP rules to completing context-dependent professional work.

PatRe patent examination mascot holding a gavel
PatRe · Patent examination

02 / Benchmarks

Research-backed evaluation,
from knowledge to procedure.

Each benchmark is presented with its original paper, code, data, and project resources.

01
Patent examination 2026 · Paper

PatRe

A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

PatRe models patent examination as an iterative process of examiner analysis and applicant response. It evaluates Office Action and rebuttal generation under direct, oracle-reference, and retrieval-simulated settings using recent USPTO cases.

480
Real-world cases
1,075
Office Actions
2
Generation tasks
02
Comprehensive IP practice ACL 2026

IPBench

Benchmarking Large Language Models on Intellectual Property Knowledge and Practice

IPBench introduces a four-level task taxonomy grounded in Depth of Knowledge. It covers technical and legal tasks across information processing, reasoning, discriminant evaluation, and creative generation in real-world IP practice.

10,374
Data instances
20
Tasks
8
IP mechanisms
03
Knowledge & consultation 2024 · Paper

IPEval

A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models

IPEval evaluates how well language models understand and apply intellectual property knowledge in agency and consultation settings. Its bilingual questions span the creation, application, protection, and management of intellectual property.

2,657
Questions
4
Capability dimensions
EN / ZH
Bilingual

03 / Scope

A structured view of
IP Intelligence.

The collection spans complementary evaluation units, task formats, and professional contexts.

Benchmark
Primary focus
Task form
Coverage
PatReProcedural
Patent examination lifecycle
Office Action and rebuttal generation
USPTO cases · multi-round histories
IPBenchComprehensive
Knowledge and real-world IP practice
Understanding, reasoning, classification, generation
English and Chinese · 8 IP mechanisms
IPEvalFoundation
IP knowledge and consultation
Multiple-choice evaluation
English and Chinese · 4 dimensions

04 / Papers

Grounded in public research papers.

Read the original work for task construction, evaluation protocols, experimental settings, findings, and limitations.

2026

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

Qiyao Wang, Xinyi Chen, Longze Chen, Hongbo Wang, Hamid Alinejad-Rokny, Yuan Lin, and Min Yang.

arXiv:2605.03571 ↗
2026

Towards IP Intelligence: Benchmarking Large Language Models on Intellectual Property Knowledge and Practice

Qiyao Wang et al. · ACL 2026.

arXiv:2504.15524 ↗
2024

IPEval: A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models

Qiyao Wang, Jianguo Huang, Shule Lu, Yuan Lin, Kan Xu, Liang Yang, and Hongfei Lin.

arXiv:2406.12386 ↗

05 / About

IP Intelligence

IPIntelligenceArena is AIforIP’s public benchmark collection for evaluating large language models across intellectual property knowledge, practice, and patent examination.

Benchmark descriptions and figures on this site are based on the corresponding papers and public repositories. Please consult the original resources for complete methodological and licensing details.

Visit AIforIP on GitHub ↗