Old Ideas, Novel Problems

The Instability of LLM-Based Novelty Evaluation

Noy Sternlicht1,2, Simra Shahid3, Peter Jansen2,4, Daniel S. Weld2,5, Pao Siangliulue2, Tom Hope1,2

1Hebrew University of Jerusalem, 2Allen Institute for AI, 3Microsoft, 4University of Arizona, 5University of Washington

Paper Code Dataset Cite

Can we trust LLM novelty judges?

The same idea pair is judged under three prompt variations: which was judged novel, which is novel, and which was judged more novel. A novelty judge scores each, and GPT-5.4 pairwise accuracy under each is 93.5, 40.9 and 59.7 on Human+Generated, versus 83.1, 76.0 and 80.5 on Human-Only.
Seemingly small changes, such as minor prompt variations, can wildly swing a judge's performance. Judges can grow markedly more brittle once LLM-generated ideas enter the data (Human+Generated), a failure that does not show up on the human-written papers they are typically tested on (Human-Only).

Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to LLMs. These judges are typically built ad hoc and validated, if at all, on human-written papers rather than on the generated ideas they are meant to score.

So, how do novelty judges perform? Not well.

We run a systematic controlled study of novelty evaluation design choices (prompt wording, retrieval, reasoning effort, idea format, and verdict aggregation), varying one at a time across six judges.

Small choices have large consequences. Simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical pairs it is shown, shifting pairwise accuracy by more than 50 points and occasionally pushing it below chance. The same change can help one judge and hurt another.

Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about the reported novelty gains of automated ideation systems, and call for more robust novelty evaluation methods.

Benchmark

Benchmark pipeline: an LLM extracts novelty signals from OpenReview submissions and reviews, filters them into validated novel and lower-novelty human-authored ideas, adds weakly labeled lower-novelty ideas from a vanilla LLM, and combines them into two data setups, Human-Only and Human+Generated.
Benchmark construction. We mine novelty labels automatically from ICLR 2026 reviews, keeping only submissions whose reviewers agree on the originality of the contribution, and complement them with weakly labeled lower-novelty ideas from a non-SOTA LLM ideator. The result is an open-source novelty evaluation dataset with two data setups, Human-Only and Human+Generated. Tap to enlarge.

Data statistics

Setup Ideas Eval. instances
D+ D− Pairwise Pointwise
Human-Only 154145ICLR 154299
Human+Generated 154154LLM 154308

D+ holds the novel ideas and D− the lower-novelty ones.

Both setups share the same D+ (154 ICLR-validated novel ideas) and differ in D−: validated lower-novelty ICLR ideas in Human-Only, LLM-generated ideas in Human+Generated.

Examples of extracted novelty signals

Praised as novel
It is the first work I have seen that attempts to constrain the diffusion process within the quotient space Quotient-Space Diffusion Models
The paper identifies and clearly defines a completely new, critical, and timely research problem Image Quality Assessment for Embodied AI
Criticized as not novel
The proposal is extremely close to NEAT (Zhong et al., 2025). Both methods replace LoRA's linear update with a nonlinear mapping… NoLoRA: Nonlinear Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
The contribution is primarily procedural integration of existing components (YOLOE + CLIP + GPT-4V + SAM2)… Consistent 3D Object Detection with Active LLM Reasoning

Controlled Study

Controlled study design: two judgment formats (pairwise, pointwise) crossed with six evaluation dimensions (judge instructions, related work access, reasoning effort, idea format, idea source, verdict aggregation) and six judge models from Anthropic and OpenAI, run one configuration at a time and reported on the Human-Only and Human+Generated setups.
Study design. Two judgment formats × six evaluation dimensions × six judges, varied one at a time and reported on both data setups. Tap to enlarge.
ChangeDescriptionProtocol
Judge promptreference: a novelty criterion defining what counts as novel
P1No guardrailsWeaker criterion, no warning against trivial combinationsboth
P2No criterionDrops the novelty definition entirelyboth
P3"Which was judged novel?"A prior review found exactly one idea novel; predict whichpairwise
P4"Which is novel?"P3 without the presumption of prior reviewpairwise
P5"Which was judged more novel?"A prior review found one idea more novel; predict whichpairwise
Evaluation setupreference: no retrieval, high reasoning, free-text abstracts, aggregated verdicts
+ RetrievalJudges also receive 5 abstracts of related workboth
Low reasoningReasoning effort lowered from high to lowboth
Idea formatIdeas rewritten as a two-field plan (purpose, mechanism)both
No aggregationA single call instead of three (pairwise: one per presentation order)both
GeneratorLower-novelty ideas regenerated by other modelsboth
Judges gpt-5.1gpt-5.2gpt-5.4 claude-sonnet-4-5claude-opus-4-5claude-opus-4-6

Results

BibTeX

@misc{sternlicht2026oldideasnovelproblems,
      title={Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation},
      author={Noy Sternlicht and Simra Shahid and Peter Jansen and Daniel S. Weld and Pao Siangliulue and Tom Hope},
      year={2026},
      eprint={2610.02022},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.02022},
}