Old Ideas, Novel Problems: The
Instability of LLM-Based Novelty Evaluation
Noy Sternlicht, Simra Shahid, Peter Jansen, Daniel S. Weld, Pao
Siangliulue, Tom Hope
Arxiv Preprint
Automated ideation systems are often evaluated on the novelty of their ideas, a judgment
increasingly delegated to LLMs. So, how do novelty judges perform? Not well. We present a systematic
controlled study of novelty evaluation design choices. We find that small changes to the prompt have
large consequences, retrieval and larger reasoning budgets help little, and purpose-built novelty
evaluators are outperformed by our cheapest prompted baseline. These results raise questions about
reported novelty gains of automated ideation systems.
LLM-as-a-Judge
Novelty Evaluation
Ideation
NLP for Science
Evaluation