default search action
Samuel Marks
Person information
Refine list
refinements active!
zoomed in on ?? of ?? records
view refined list in
2020 – today
- 2026
- [i31]Abhay Sheshadri, Aidan Ewart, Kai Fronsdal, Isha Gupta, Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang:
AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors. CoRR abs/2602.22755 (2026) - [i30]Helena Casademunt, Bartosz Cywinski, Khoi Tran, Arya Jakkli, Samuel Marks, Neel Nanda:
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation. CoRR abs/2603.05494 (2026) - [i29]James Chua, Jan Betley, Samuel Marks, Owain Evans:
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious. CoRR abs/2604.13051 (2026) - [i28]Keshav Shenoy, Li Yang, Abhay Sheshadri, Sören Mindermann, Jack Lindsey, Samuel Marks, Rowan Wang:
Introspection Adapters: Training LLMs to Report Their Learned Behaviors. CoRR abs/2604.16812 (2026) - [i27]Chloe Li, Nevan Wichers, Sara Price, Samuel Marks, Jonathan Kutasov:
Model Spec Midtraining: Improving How Alignment Training Generalizes. CoRR abs/2605.02087 (2026) - [i26]Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks:
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens. CoRR abs/2607.03502 (2026) - 2025
- [c6]Jaden Fried Fiotto-Kaufman, Alexander Russell Loftus, Eric Todd, Jannik Brinkmann, Koyena Pal, Dmitrii Troitskii, Michael Ripa, Adam Belfki, Can Rager, Caden Juang, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Nikhil Prakash, Carla E. Brodley, Arjun Guha, Jonathan Bell, Byron C. Wallace, David Bau:
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals. ICLR 2025 - [c5]Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller:
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. ICLR 2025 - [c4]Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Demian Till, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda:
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. ICML 2025 - [c3]Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau:
Erasing Conceptual Knowledge from Language Models. NeurIPS 2025 - [i25]Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Isaac Bloom, David Chanin
, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, Matthew Wearden, Arthur Conmy, Samuel Marks, Neel Nanda:
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. CoRR abs/2503.09532 (2025) - [i24]Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra-Sharma, Daniel M. Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R. Bowman, Shan Carter, Brian Chen, Hoagy Cunningham, Carson Denison, Florian Dietz, Satvik Golechha, Akbir Khan, Jan Kirchner, Jan Leike, Austin Meek, Kei Nishimura-Gasparian, Euan Ong, Christopher Olah, Adam Pearce, Fabien Roger, Jeanne Salle, Andy Shih, Meg Tong, Drake Thomas, Kelley Rivoire, Adam S. Jermyn, Monte MacDiarmid, Tom Henighan, Evan Hubinger:
Auditing language models for hidden objectives. CoRR abs/2503.10965 (2025) - [i23]Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase, Samuel Marks, Jacob Goldman-Wetzler, Linda Petrini, Henry Sleight, Collin Burns, He He, Shi Feng, Ethan Perez, Jan Leike:
Unsupervised Elicitation of Language Models. CoRR abs/2506.10139 (2025) - [i22]Adam Karvonen, Samuel Marks:
Robustly Improving LLM Fairness in Realistic Settings via Interpretability. CoRR abs/2506.10922 (2025) - [i21]Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, Owain Evans:
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data. CoRR abs/2507.14805 (2025) - [i20]Helena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks, Senthooran Rajamanoharan, Neel Nanda:
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning. CoRR abs/2507.16795 (2025) - [i19]Bartosz Cywinski, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan, Neel Nanda, Arthur Conmy, Samuel Marks:
Eliciting Secret Knowledge from Language Models. CoRR abs/2510.01070 (2025) - [i18]Nevan Wichers, Aram Ebtekar, Ariana Azarbal, Victor Gillioz, Christine Ye, Emil Ryd, Neil Rathi, Henry Sleight, Alex Mallen, Fabien Roger, Samuel Marks:
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment. CoRR abs/2510.05024 (2025) - [i17]Stewart Slocum, Julian Minder, Clément Dumas, Henry Sleight, Ryan Greenblatt, Samuel Marks, Rowan Wang:
Believe It or Not: How Deeply do LLMs Believe Implanted Facts? CoRR abs/2510.17941 (2025) - [i16]Tim Tian Hua, Andrew Qin, Samuel Marks, Neel Nanda:
Steering Evaluation-Aware Language Models to Act Like They Are Deployed. CoRR abs/2510.20487 (2025) - [i15]Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks:
Liars' Bench: Evaluating Lie Detectors for Language Models. CoRR abs/2511.16035 (2025) - [i14]Ching Fang, Samuel Marks:
Unsupervised decoding of encoded reasoning using language model interpretability. CoRR abs/2512.01222 (2025) - [i13]Jordan Taylor, Sid Black, Dillon Bowen, Thomas Read
, Satvik Golechha, Alex Zelenka-Martin, Oliver Makins, Connor Kissane, Kola Ayonrinde, Jacob Merizian, Samuel Marks, Chris Cundy, Joseph Isaac Bloom:
Auditing Games for Sandbagging. CoRR abs/2512.07810 (2025) - [i12]Adam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente, Subhash Kantamneni, Julian Minder, Euan Ong, Arnab Sen Sharma, Daniel Wen, Owain Evans, Samuel Marks:
Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers. CoRR abs/2512.15674 (2025) - 2024
- [c2]Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks:
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models. NeurIPS 2024 - [c1]Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger B. Grosse, Owain Evans:
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data. NeurIPS 2024 - [i11]Samuel Marks, Can Rager
, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller
:
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. CoRR abs/2403.19647 (2024) - [i10]Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, Evan Hubinger:
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. CoRR abs/2406.10162 (2024) - [i9]Johannes Treutlein, Dami Choi, Jan Betley
, Cem Anil, Samuel Marks, Roger Baker Grosse, Owain Evans:
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data. CoRR abs/2406.14546 (2024) - [i8]Jaden Fiotto-Kaufman, Alexander R. Loftus, Eric Todd
, Jannik Brinkmann, Caden Juang, Koyena Pal, Can Rager, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Michael Ripa, Adam Belfki, Nikhil Prakash, Sumeet Multani, Carla E. Brodley, Arjun Guha, Jonathan Bell, Byron C. Wallace, David Bau:
NNsight and NDIF: Democratizing Access to Foundation Model Internals. CoRR abs/2407.14561 (2024) - [i7]Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, Samuel Marks:
Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models. CoRR abs/2408.00113 (2024) - [i6]Aaron Mueller
, Jannik Brinkmann, Millicent L. Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, Eric Todd
, David Bau, Yonatan Belinkov:
The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability. CoRR abs/2408.01416 (2024) - [i5]Rohit Gandikota, Sheridan Feucht, Samuel Marks, David Bau:
Erasing Conceptual Knowledge from Language Models. CoRR abs/2410.02760 (2024) - [i4]Adam Karvonen, Can Rager, Samuel Marks, Neel Nanda:
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks. CoRR abs/2411.18895 (2024) - [i3]Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Samuel Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, Sören Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Evan Hubinger:
Alignment faking in large language models. CoRR abs/2412.14093 (2024) - 2023
- [j1]Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphaël Ségerie, Micah Carroll, Andi Peng, Phillip J. K. Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca D. Dragan, David Krueger, Dorsa Sadigh, Dylan Hadfield-Menell:
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Trans. Mach. Learn. Res. 2023 (2023) - [i2]Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphaël Ségerie, Micah Carroll, Andi Peng, Phillip J. K. Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca D. Dragan, David Krueger, Dorsa Sadigh, Dylan Hadfield-Menell:
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. CoRR abs/2307.15217 (2023) - [i1]Samuel Marks, Max Tegmark:
The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. CoRR abs/2310.06824 (2023)
Coauthor Index
manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.
Unpaywalled article links
Add open access links from to the list of external document links (if available).
Privacy notice: By enabling the option above, your browser will contact the API of unpaywall.org to load hyperlinks to open access articles. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Unpaywall privacy policy.
Archived links via Wayback Machine
For web page which are no longer available, try to retrieve content from the of the Internet Archive (if available).
Privacy notice: By enabling the option above, your browser will contact the API of archive.org to check for archived content of web pages that are no longer available. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Internet Archive privacy policy.
Reference lists
Add a list of references from ,
, and
to record detail pages.
load references from crossref.org and opencitations.net
Privacy notice: By enabling the option above, your browser will contact the APIs of crossref.org, opencitations.net, and semanticscholar.org to load article reference information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the Crossref privacy policy and the OpenCitations privacy policy, as well as the AI2 Privacy Policy covering Semantic Scholar.
Citation data
Add a list of citing articles from and
to record detail pages.
load citations from opencitations.net
Privacy notice: By enabling the option above, your browser will contact the API of opencitations.net and semanticscholar.org to load citation information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the OpenCitations privacy policy as well as the AI2 Privacy Policy covering Semantic Scholar.
OpenAlex data
Load additional information about publications from .
Privacy notice: By enabling the option above, your browser will contact the API of openalex.org to load additional information. Although we do not have any reason to believe that your call will be tracked, we do not have any control over how the remote server uses your data. So please proceed with care and consider checking the information given by OpenAlex.
last updated on 2026-08-18 00:18 CEST by the dblp team
all metadata released as open data under CC0 1.0 license
see also: Terms of Use | Privacy Policy | Imprint