Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–3 of 3 results for author: Gronas, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.01152  [pdf, ps, other

    cs.CL

    AGC-Bench: Measuring Artificial General Creativity

    Authors: Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai, Anna Attuch, Namrata Shivagunde, Swastik Roy, Rajkumar Pujari, Paul V. DiStefano, Sherin Muckatira, Claire E. Stevenson, Mikhail Gronas, Anna Rumshisky

    Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both questions now apply to LLMs, but a unified benchmark of AI creativity remains elusive. We introduce AGC-Bench, an artificial general creativity benchmark built from a systematic review of the AI creativity literature (3,101… ▽ More

    Submitted 1 July, 2026; v1 submitted 1 July, 2026; originally announced July 2026.

  2. arXiv:2606.06526  [pdf, ps, other

    cs.AI cs.LG

    CrowdMath: A Dataset of Crowdsourced Mathematical Research Discussions

    Authors: Sherin Muckatira, Jesse Geneson, Slava Gerovitch, Pavel Etingof, Mikhail Gronas, Anna Rumshisky

    Abstract: Large language models have made substantial progress on mathematical reasoning, but existing benchmarks typically evaluate well-specified problems with final answers, step-by-step solutions, or complete proofs. They do not capture collaborative open-problem solving: a setting in which participants propose partial arguments, identify gaps or errors in prior steps, repair flawed reasoning, and gradu… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 16 pages, 4 figures

  3. arXiv:2605.27832  [pdf, ps, other

    cs.CL

    Playing with Words, Improving with Rewards: Training Language Models for Creative Association

    Authors: Vijeta Deshpande, Namrata Shivagunde, Sherin Muckatira, Hadrien Glaude, Mikhail Gronas, Claire Stevenson, Roger Beaty, Anna Rumshisky

    Abstract: Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet the subjective nature of creativity and the limits of human judgment make training LLMs for creativity especially challenging. As a solution, we train LLMs on Codenames, a word-association game that exercises the two cent… ▽ More

    Submitted 26 May, 2026; originally announced May 2026.