Brief
Judge-guided revision narrows the gap between cheap and costly patent drafts
A preprint scores AI-drafted patents with a separate model acting as judge, then feeds that feedback back to the drafting agent. Judge-guided revision improved judge-assessed quality and let a low-reasoning agent approach a much more expensive one, but the same judge agreed with a professional patent attorney only in metric-dependent ways.
The paper introduces Vibe Patenting, an end-to-end testbed where a separately invoked LLM judge evaluates generated patent drafts and provides structured feedback for revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improved judge-assessed quality, while unguided revision tended to saturate.
The useful detail: iterative judge feedback allowed a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improved quality, and domain-specific agentic workflows added further gains. However, validation against a professional patent attorney found meaningful but strongly metric-dependent agreement and systematic calibration differences.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by arXiv
- Judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate.
- Iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent.
- The judge was validated against independent evaluation by a professional patent attorney, finding meaningful but strongly metric-dependent agreement and systematic calibration differences.
- Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains.
Sources
- arXivText stored 15 September 2026
How this story was checked. Written from the 1 page listed above, stored 15 September 2026; claims checked against that stored text on 15 September 2026.
What that means
- 4 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.