Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to LLMs. These judges are typically built ad hoc and validated, if at all, on human-written papers rather than on the generated ideas they are meant to score.
So, how do novelty judges perform? Not well.
We run a systematic controlled study of novelty evaluation design choices (prompt wording, retrieval, reasoning effort, idea format, and verdict aggregation), varying one at a time across six judges.
Small choices have large consequences. Simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical pairs it is shown, shifting pairwise accuracy by more than 50 points and occasionally pushing it below chance. The same change can help one judge and hurt another.
Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about the reported novelty gains of automated ideation systems, and call for more robust novelty evaluation methods.
| Setup | Ideas | Eval. instances | ||
|---|---|---|---|---|
| D+ | D− | Pairwise | Pointwise | |
| Human-Only | 154 | 145ICLR | 154 | 299 |
| Human+ |
154 | 154LLM | 154 | 308 |
D+ holds the novel ideas and D− the lower-novelty ones.
Both setups share the same D+ (154 ICLR-validated novel ideas) and differ in D−: validated lower-novelty ICLR ideas in Human-Only, LLM-generated ideas in Human+Generated.
It is the first work I have seen that attempts to constrain the diffusion process within the quotient spaceQuotient-Space Diffusion Models
The paper identifies and clearly defines a completely new, critical, and timely research problemImage Quality Assessment for Embodied AI
The proposal is extremely close to NEAT (Zhong et al., 2025). Both methods replace LoRA's linear update with a nonlinear mapping…NoLoRA: Nonlinear Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
The contribution is primarily procedural integration of existing components (YOLOE + CLIP + GPT-4V + SAM2)…Consistent 3D Object Detection with Active LLM Reasoning
| Change | Description | Protocol |
|---|---|---|
| Judge promptreference: a novelty criterion defining what counts as novel | ||
| P1No guardrails | Weaker criterion, no warning against trivial combinations | both |
| P2No criterion | Drops the novelty definition entirely | both |
| P3"Which was judged novel?" | A prior review found exactly one idea novel; predict which | pairwise |
| P4"Which is novel?" | P3 without the presumption of prior review | pairwise |
| P5"Which was judged more novel?" | A prior review found one idea more novel; predict which | pairwise |
| Evaluation setupreference: no retrieval, high reasoning, free-text abstracts, aggregated verdicts | ||
| + Retrieval | Judges also receive 5 abstracts of related work | both |
| Low reasoning | Reasoning effort lowered from high to low | both |
| Idea format | Ideas rewritten as a two-field plan (purpose, mechanism) | both |
| No aggregation | A single call instead of three (pairwise: one per presentation order) | both |
| Generator | Lower-novelty ideas regenerated by other models | both |
On pairwise Human+Generated, asking which idea a prior review judged novel (P3) improves every judge, by more than 20 points in some cases. Yet P4, a close variant of the same prompt that makes no reference to prior human judgment, makes gpt-5.4 collapse below chance. Meanwhile, the fixes prior work commonly reaches for have a limited effect: lowering reasoning effort from high to low costs at most 7 points and is usually neutral, and retrieval improves only some judges, by at most 16 points.
The same change might be neutral or even beneficial for one judge, but destructive for another. This is seen most clearly in pointwise judging: removing the guardrails from the novelty definition (P1) adds 15.7 points for claude-sonnet-4-5 on Human-Only, yet costs claude-opus-4-6 6.7. A configuration validated on one judge therefore carries no guarantee for the next. In line with prior work, pairwise judging is generally more stable than pointwise judging.
claude-sonnet-4-5 exposes no reasoning_effort parameter, hence the blank entry. Tap to enlarge.A pairwise judge ties when its aggregated verdicts favor neither idea. Depending on the configuration, even relatively strong judges like claude-opus-4-6 can tie on 35% of the pairs, and weaker ones like claude-sonnet-4-5 on more than half. Every judge ties more often on Human+Generated than on Human-Only, roughly twice as often for claude-opus-4-5 and claude-sonnet-4-5.
We replace our retrieval pipeline with a strong agentic RAG judge: gpt-5.6-sol at xhigh reasoning effort, searching the literature iteratively with no cap on searches. It reaches 0.78 macro-F1, 11 points behind claude-opus-4-6 with the simple five-abstract pipeline, and within noise of claude-opus-4-5 and gpt-5.4, despite costing 10× more per decision.
We compare prompted judges to two dedicated novelty judges, Idea Novelty Checker (Shahid et al., 2025) and AI-Scientist (Yamada et al., 2025). In every setting and under both backbones (claude-opus-4-6 and gpt-5.4), our cheapest prompted configuration outscores both by 8 to 29 macro-F1 points, while costing up to 32× less. The Pareto front consists entirely of simple prompted judges.
claude-opus-4-6, on Human-Only (left) and Human+Generated (right). The Pareto front (line) contains only prompted judges. Error bars: 95% CIs from a paired bootstrap. Tap to enlarge.@misc{sternlicht2026oldideasnovelproblems,
title={Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation},
author={Noy Sternlicht and Simra Shahid and Peter Jansen and Daniel S. Weld and Pao Siangliulue and Tom Hope},
year={2026},
eprint={2610.02022},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2610.02022},
}