Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–1 of 1 results for author: Luong, I

Searching in archive cs. Search in all archives.
.
  1. arXiv:2605.16301  [pdf, ps, other

    cs.CY cs.AI cs.LG

    Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

    Authors: Isabella Luong, Joyee Chen, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu

    Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks… ▽ More

    Submitted 31 July, 2026; v1 submitted 18 April, 2026; originally announced May 2026.