Computers and Society
See recent articles
Showing new listings for Friday, 21 August 2026
- [1] arXiv:2608.19278 [pdf, other]
-
Title: Mapping General-Purpose AI Governance in Twenty AI Middle-Power JurisdictionsComments: for associated dataset file, see doi:https://doi.org/10.5281/zenodo.21978946Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
The most capable general-purpose AI (GPAI) models are mostly built in two jurisdictions, the United States and China, but the risks they carry land globally. Regionally advanced economies hosting no frontier developer, which we call AI middle-powers, are writing their own rules to govern GPAI. This paper investigates which GPAI-relevant provisions these AI middle-powers have enacted, mapping twenty jurisdictions including the European Union at the level of the individual provision, across four governance areas that trace the accountability chain for the model layer: systemic risk assessment, evaluation and verification, prohibitions with monitoring and detection, and serious incident reporting. Confirmed absence is recorded as data alongside positive provision. We find that jurisdictions converge on form, but diverge on force. Sixteen engage in at least three of the four governance areas, yet only about one in five provisions sit in binding law, and three-quarters of the instruments that do bind do so without defining GPAI. The institutional infrastructure shows the same shape: four in five of the mapped governance actors hold mandates that predate GPAI, and obligations attach wherever the inherited regime already reached, which is the application layer rather than the model. Where these states engage the model layer, they build capacity to observe it rather than impose duties on those who build it, and almost every evaluation body was constituted without the power to act on what it finds. Nominal coverage of the full accountability chain reaches eleven jurisdictions, but only five hold more than one provision in every area and, outside the EU, no jurisdiction imposes a binding evaluation duty on a model developer. The dataset gives researchers and policymakers a provision-level basis for identifying where regimes could align, and where coordination would have to start from scratch.
- [2] arXiv:2608.19379 [pdf, other]
-
Title: Multi-Tier Mentorship with AI-Assisted Development: Authentic Engineering for K-12 and UndergraduatesComments: 8 pages, 8 figures, 2 tables. Accepted to IEEE ISEC 2026Subjects: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
K-12 students often possess creative engineering ideas but lack technical skills to build them, while undergraduates have coding expertise but few opportunities to lead real-world projects or mentor others. The rapid development of AI-assisted tools offers a potential bridge to connect these groups, yet the structure for effective K-12 and university collaborations remains underexplored. This paper introduces a multi-tiered mentorship framework enabling high school students to engage in authentic engineering through AI-assisted development using large language models and AI agents, while undergraduate mentors provide architectural oversight. We test this framework through LuckyTag, a privacy-preserving NFC-based lost-and-found system. The model positions high schoolers as product leads, undergraduates as technical architects, and faculty as minimal-intervention advisors. A pilot with four high school students, three undergraduates and two faculty yielded survey data showing high perceived barrier removal and gains in system architecture understanding. Thematic analysis reveals that AI amplifies rather than supplants mentoring demands, requiring human oversight for logic and security. These findings suggest a hybrid model for equitable K-12 and university collaboration on computing integration that emphasizes "AI micromanagement" and architectural reasoning over traditional syntax.
- [3] arXiv:2608.19390 [pdf, html, other]
-
Title: Navigating Epistemic Monocultures in AI-Driven Science: A Simulation StudyComments: Accepted at Philosophy of ScienceSubjects: Computers and Society (cs.CY)
AI integration into scientific communities promises accelerated discovery but raises concerns about detrimental homogenization. We develop an NK landscape model to explore these promises and risks. We find that non-personalized AI systems that offer uniform guidance yield benefits only under a narrow conjunction of problem structure, practices, and baseline research capabilities, becoming harmful otherwise. We implement two proposed mitigations: randomization and personalization. While randomization's utility remains restricted to decomposable problems, personalization can enhance diversity, enabling benefits across a broader range of conditions. Crucially, these benefits are not automatic, but depend on effective institutional adaptation, requiring new standards and practices.
- [4] arXiv:2608.19495 [pdf, html, other]
-
Title: A Scoping Review of Methods to Measure the Energy and Carbon Footprint of Web Tracking and AdvertisingComments: 5 pages, 0 figures. Accepted at the 2nd International Workshop on Low Carbon Computing (LOCO 2026), Lancaster University, United Kingdom, 10-11 September 2026. Part of the LOCO 2026 proceedings, arXiv:2608.02072Subjects: Computers and Society (cs.CY)
The environmental impact of web tracking and advertising is increasingly receiving attention as the ICT sector's carbon footprint keeps rising. Yet the scholarship addressing this question remains scattered across disciplines and inconsistent in its terminology. This paper presents a scoping review of the literature on methods for measuring the energy and carbon footprint of web tracking and advertising. From an initial pool of 46 articles identified through a structured title-based search on Google Scholar, we arrived at a final corpus of 15 papers, from which we identified five distinct methodological approaches: ad blocking, controlled environment, replaying ads, traffic flow analysis, and literature-derived estimation. This review provides a structured overview of the current methodological landscape and a foundation for more comprehensive environmental accounting of the ad tech ecosystem.
- [5] arXiv:2608.19545 [pdf, other]
-
Title: Two-sided receptivity to conversational AI agents in online dating: Bilingual survey data from Fledge.LoveSubjects: Computers and Society (cs.CY); Information Retrieval (cs.IR)
Autonomous conversational agents and generative-AI features are being added to online dating platforms faster than public evidence about user attitudes can accumulate, and the scarcest evidence concerns the receiving side: how people react when the profiles, messages, or conversation partners they encounter are machine-generated. We release two anonymized survey datasets collected from active users of this http URL, a dating platform serving an international user base. The first (N = 2,617; Russian and English forms) measures receptivity to autonomous conversational agents with a seven-item battery that separates the principal role (deploying one's own agent) from the counterpart role (encountering someone else's), plus six ordinal covariates and two auxiliary items. The second (N = 2,894) measures interest in three passive generative-AI features. The release includes model-derived scores for 2,499 complete cases, a bilingual codebook, a documented anonymization pipeline with a k-anonymity audit, executable analysis notebooks, and canonical outputs, supporting reuse in human-AI communication, recommender-systems, and cross-cultural technology-acceptance research.
- [6] arXiv:2608.19616 [pdf, html, other]
-
Title: Modeling AI Overreliance as a Complex Adaptive SystemComments: Accepted in CSS 2026Subjects: Computers and Society (cs.CY)
Whether AI assistance helps or harms a population depends less on the model's accuracy than on whether people rely on it appropriately trusting it when it is right and checking it when it is not. Yet reliance is usually studied one user at a time. We model it as a population process: agents repeatedly solve a task alone, accept an AI answer, or verify it, updating a Bayesian belief about AI quality and, when networked, learning from peers. Four results form one story. The environment sets the baseline: task difficulty and AI quality fix both overreliance and calibration regret. Social learning creates consensus, not overreliance: a mean-preservation theorem, confirmed by a 2*2 topology*tagging design, shows connectivity moves the aggregate only when influence transmits beliefs. Social proof turns reliance into a feedback cascade: visible unverified use suppresses verification and tips the population into collective overreliance. Feedback design can prevent collapse: making verification visible or dampening social proof reverses it. Together, the results frame AI reliance as a computational social dynamics problem, where individual learning, peer observation, and feedback exposure jointly shape whether a population remains calibrated.
- [7] arXiv:2608.19707 [pdf, html, other]
-
Title: ChatGPT Solves All Tested Qiskit Homework AssignmentsComments: 5 pagesSubjects: Computers and Society (cs.CY); Quantum Physics (quant-ph)
Generative AI creates an assessment challenge in quantum software education: a student can provide a homework notebook to ChatGPT and request a completed submission. This study examined whether introductory Qiskit homework could remain autogradable while requiring students to run, review, and discuss results rather than banning AI. Three packages were tested: seeded basis-state circuits with bit flips and customized measurement mappings; Quantum Fourier Transform followed by inverse-transform recovery; and seeded Deutsch-Jozsa with customized oracle masks. The designs used personalization, simulator execution, JSON submissions, hidden references, circuit metrics, reflections, and optional IBM Quantum execution. For each package, one student-visible instance was tested in 50 separate ChatGPT sessions, yielding 150 sessions overall. Every final artifact was executed and passed its grader. Nine sessions were fully archived; none required operator code changes or correction of quantum logic. Under the study's operational definition, each tested instance had zero observed ChatGPT-resiliency. Seeds changed parameters rather than task structure, expected results remained derivable from visible assignment logic, scaffolding exposed key solution steps, and hidden grading verified output consistency without establishing independent authorship or understanding. Because one instance was repeated for each package, the results do not establish solvability for every seed or possible Qiskit assessment. The tested personalized, execution-oriented take-home designs therefore did not prevent successful completion under a minimally engaged-student workflow. Correct artifacts should be complemented by direct assessment through supervised modification, oral defense, prediction, and transfer tasks.
- [8] arXiv:2608.19816 [pdf, html, other]
-
Title: Understanding as an Explicit and Assessable Component of Frontier AI Safety DecisionsStephen Barrett, Robin Bloomfield, Alexandra Chirilă, Mamoon Masud, David Meredith Hardy, Phillip MulvanaSubjects: Computers and Society (cs.CY)
Decision makers need sufficient understanding to make good decisions about complex AI systems. However, AI deployment decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgin and Arendt. To assess the methodology we trialled two different scenarios. One scenario, which we investigated through role-based analysis, concerned the risk of scheming in the deployment of an AI coding agent in a robotics company and the other scenario was for the higher uncertainty, more decision-critical argument of 'If Anyone Builds It, Everyone Dies' (Yudkowsky and Soares). The trial's central finding, for these two scenarios, is that the methodology could be applied and was found to be generative: we found the analyses that justify sufficiency of understanding (internal coherence, tethering, felicitous falsehoods, external coherence) drives the engineering.
New submissions (showing 8 of 8 entries)
- [9] arXiv:2608.19216 (cross-list from cs.AI) [pdf, html, other]
-
Title: Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelComments: 25 pages, 5 figures. Synthetic access-ablation study; not real-world payment-system evidenceSubjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs. This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through contracts, architecture, audit, vendor assurance, residual risk, or reduced system scope. It also reports a synthetic access-ablation experiment over 1.35 million synthetic case simulations and interprets the findings through an anonymised national-payments-infrastructure scenario. The experiment is not real-world payment-system evidence; it is a construct-validity exercise. The results show that complete logs improve diagnosis, a pre-execution gateway enables intervention, trace access and model-version control strengthen post-incident explanation, and scope restriction can improve safety while reducing usefulness. Control protocols proposed as general safety solutions should therefore state their access assumptions explicitly.
- [10] arXiv:2608.19224 (cross-list from stat.ME) [pdf, html, other]
-
Title: Causal Inference under Interference with Learned Exposure MappingsComments: 20 pages, 6 figuresSubjects: Methodology (stat.ME); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Applications (stat.AP)
Exposure mappings are often assumed to be known in causal spillover analyses. In environmental settings, however, they are typically induced by transport processes that are not directly observed and must instead be learned from pollution data. We study how uncertainty in learned transport processes propagates into exposure mappings and downstream spillover inference under interference. We compare mechanistic transport models with modern operator-learning approaches, including PDE, PINO, FNO, and GeoPT, using both simulation studies and an empirical analysis of California PM$_{2.5}$ data. In simulations, all four transport models achieved nearly identical pollution prediction accuracy, yet estimated spillover effects ranged from 1.78 to 2.27. Models that more accurately recovered the induced exposure mapping also produced spillover estimates closer to the true effect. Disagreement was modest for regional interventions but substantially larger for localized point-source interventions. The California analysis showed the same pattern: competing transport models produced similar predictions of observed PM${2.5}$ concentrations while implying different spillover effects under hypothetical pollution-control interventions. Our findings suggest that predictive agreement alone is insufficient for reliable causal inference when exposure mappings are learned rather than directly observed.
- [11] arXiv:2608.19433 (cross-list from cs.SI) [pdf, html, other]
-
Title: Social.Wiki: A Web Held in CommonTheia Henderson, Carmel Schare, Ana Dodik, Clemens N. Klokmose, Ziv Epstein, David D. Clark, David R. KargerComments: Accepted to The 39th Annual ACM Symposium on User Interface Software and Technology (UIST '26), November 02-05, 2026, Detroit, MI, USASubjects: Social and Information Networks (cs.SI); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
Many of the websites people depend on have owners whose interests are not fully aligned with their users. We address the root of this problem by presenting a reimagining of the web where sites are not owned at all but are instead collaboratively produced like Wikipedia articles. We call the system this http URL because it supports the co-creation of interactive social sites, such as those for microblogging, messaging, dating, gaming, ride sharing, and so on. With off-the-shelf AI tools, people with little or no programming experience can edit these sites to better reflect the needs and preferences of their communities.
this http URL builds on ideas from collaborative malleable software systems such as Webstrates, but is designed for public participation rather than use only within small, trusted groups. To this end, this http URL includes governance to mitigate conflict. To accommodate diverse governance preferences, our model of "plural governance" lets people independently choose the policies that determine which edits to a site they see. this http URL also implements a granular security model to protect personal data in a malleable environment.
Complementing the decentralized design and governance of this http URL sites, both site edits and within-site data are stored on Graffiti, a decentralized infrastructure, decoupling the ownership of underlying servers from the ownership of sites.
We evaluate this http URL through case studies that demonstrate the range of sociotechnical structures it supports, as well as through deployments at a hackathon and in the wild. - [12] arXiv:2608.19437 (cross-list from cs.CL) [pdf, html, other]
-
Title: Are LLMs becoming similarly creative? Evidence from three years of modelsComments: 12 pages, 4 figuresSubjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.
- [13] arXiv:2608.19922 (cross-list from cs.LG) [pdf, html, other]
-
Title: Auditing Recorded Predictive Lead Service-Line Classifications Against Physical Verification: A Statewide Study of New YorkComments: 14 pages, 3 figures, 3 tables. Analysis code and data provenance: this https URLSubjects: Machine Learning (cs.LG); Computers and Society (cs.CY)
Under the US Lead and Copper Rule Revisions, a utility may determine a service line's material with a predictive model instead of inspecting it. New York State publishes, per address, which method was used. Almost no address carries both a model classification and a physical verification, so the check is between populations within a utility rather than paired addresses. We screen all 153 New York localities that classified at least 100 addresses this way. Seventy-five (49%), covering 125,990 addresses or 57% of those screened, record one value. Zero variance alone is not misconduct: 68 of the 75 match their own verification or have too little to test. Seven are contradicted by their own crews, six beyond any sampling explanation. Five are boroughs of New York City, which file as one system; one is East Rochester, 550 km away. New York City is the largest case: a predictive model is the recorded basis for 43,215 addresses, and on all of them the recorded material is "Known Other". The city records "Unknown" on 121,779 addresses, 1,880 already excavated, and lead on 120,692. In the model bucket both counts are zero, and the 95% upper bound on the rate is 0.0085%. Across the rest of New York the same method records lead or the hedge "Unknown but could be lead" on 12.21% of 176,888 addresses, a comparison whose weaknesses we report. The model-cleared population is newer, median year built 1984 against 1930, and construction era accounts for about a third of the gap and not the rest: holding era fixed, records-based classification finds lead at 4.3-31.9%, physical verification at 1.5-14.5%, the model in no era. Six era-aware estimators place the expected lead lines among them at 1,150-1,450. Two findings need no comparison: 7,782 of these addresses are in pre-1940 buildings, and the archived 2025 snapshot shows the public-side determination was copied from a customer-side model output.
- [14] arXiv:2608.19950 (cross-list from cs.HC) [pdf, html, other]
-
Title: Designing Human-mediated AI Guidance: Ready Together for Personalized Family Emergency PreparednessComments: Accepted for Italian Workshop on Artificial Intelligence for Human-Machine Interaction (AIxHMI 2026), October 6-9, 2026, Perugia, ItalySubjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Artificial intelligence (AI) systems are increasingly used across domains to provide personalized information, recommendations, and decision support. However, in some contexts, AI-generated information may not be suitable for direct delivery to the final recipient. Instead, it may need to be interpreted, adapted, and communicated by a human who understands the recipient's needs, emotional state, and situational context. Human-AI interaction research has given less attention to situations in which a more knowledgeable human acts as an intermediary between an AI system and a less experienced or less informed recipient. We introduce the human-mediated AI guidance framework and explore it through Ready Together, an AI-supported family emergency preparedness system in which parents mediate AI-generated content for their children. The system is designed to provide personalized guidance and support parents in making emergency preparedness more interactive and understandable through guided activities and family-centered learning. The system design was informed by a qualitative, design-oriented research process involving semi-structured interviews and co-design activities. Findings identified challenges in family emergency preparedness, including difficulty discussing emergencies with children, uncertainty about providing appropriate explanations, and a preference for interactive learning activities. These findings informed the design of an interactive prototype, subsequently evaluated through a pilot study and a heuristic evaluation. Participants responded positively to the personalized recommendations and practical activities. Preliminary findings suggest that human-mediated AI guidance may support context-sensitive family preparedness while preserving parents' responsibility for interpreting, adapting, and communicating AI-generated information.
- [15] arXiv:2608.20041 (cross-list from cs.AI) [pdf, html, other]
-
Title: A three-dimensional typology of agency for advanced AI systemsSubjects: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions in philosophy, ethics, legal theory and sociology, we develop a typology of agency for frontier AI systems consisting of three dimensions: the nature of agency (moral or legal), its mode (individual or collective) and its locus (human or non-human). Combining these dimensions produces eight possible instantiations of agency, which we classify as conventional, contested or controversial. The typology separates legal from moral agency and thereby creates conceptual space for considering individual, legal, non-human agency without presupposing that advanced AI systems are moral agents. We argue that this distinction is increasingly relevant where instrumental goal pursuit complicates the attribution of AI actions to particular human actors.
- [16] arXiv:2608.20145 (cross-list from cs.CR) [pdf, other]
-
Title: Trustworthy mobile edge caching: a blockchain approach to mitigate malicious nodes and incentivize cache sharingSubjects: Cryptography and Security (cs.CR); Computers and Society (cs.CY)
As mobile network traffic continues to grow, content caching on edge servers is critical for reducing latency. However, challenges such as malicious edge servers that may delete or manipulate cached content, along with the limited capacity of these servers, need to be addressed. To overcome the capacity limitations, helper mobile nodes can contribute their cache resources. However, due to their selfish behavior, an incentive mechanism is necessary to encourage resource sharing. Additionally, these helper nodes can also be malicious. This paper proposes a blockchain-based trust management mechanism that addresses these challenges by accurately identifying trustworthy edge servers and mobile nodes. The proposed mechanism calculates both direct and indirect trust using smart contracts, ensuring that malicious nodes are effectively filtered out. Trustworthiness is determined based on mobile node satisfaction with the quality of service, and trust data is securely stored on the blockchain. To combat node selfishness, a reward mechanism is introduced to incentivize cache sharing. Furthermore, a blockchain-based authentication mechanism protects against node impersonation. Our approach optimizes trust, cache capacity, and cost efficiency while considering mobile node mobility, energy consumption, and computational power constraints during the consensus process. Simulation results show that the proposed method can accurately distinguish between honest and malicious servers, even with a 10% noise in data.
- [17] arXiv:2608.20160 (cross-list from cs.CR) [pdf, other]
-
Title: Chameleon: Robust Defense Against Tor Website Fingerprinting via Many-to-Many Traffic MorphingSubjects: Cryptography and Security (cs.CR); Computers and Society (cs.CY)
Website fingerprinting (WF) attacks can infer users' browsing activities from encrypted Tor traffic by exploiting side-channel features. Although many WF defenses have been proposed, we find that most existing defenses create learnable web trace mapping features. We further show that robustness against adversarial training does not necessarily imply robustness against defense-aware autoencoder (DAAE)-based attacks.
To address these limitations, we present Chameleon, a robust WF defense based on many-to-many randomized traffic morphing. Chameleon selects morphing candidates with high intra-class diversity and low inter-class disparity. Chameleon randomly maps each webpage trace to multiple candidates, and allows different webpages to share morphing targets, thereby increasing adversarial uncertainty. For practical Tor deployment, Chameleon introduces a radix-trie-based synchronization mechanism that enables pluggable transport (PT) endpoints to identify consistent morphing traces using packet-direction prefixes, together with trace mutation and normalized prefix matching to reduce overhead. We evaluate Chameleon against six state-of-the-art defenses and five WF attacks on three public datasets in closed- and open-world settings. Compared with Adaptive Tamaraw, Chameleon reduces adversarial-training-based attack accuracy by up to 36.74% while reducing bandwidth and time overhead by 34.12% and 60.38%, respectively. Under DAAE-based RF attacks on GTT23, Chameleon limits attack performance to 35.19% F1-score while Adaptive Tamaraw only limits it to 88.22% F1-score. In the real-world PT bridge evaluation, Chameleon substantially reduces the effectiveness of strong WF attacks while incurring only 16.25% time overhead. - [18] arXiv:2608.20202 (cross-list from cs.AI) [pdf, html, other]
-
Title: MemTrapBench: Benchmarking Cognitive Traps in LLM Memory UseMengru Wang, Haozhe Luo, Zhenqian Xu, Zhixiang Cui, Haoming Xu, Qu Yang, Jizhan Fang, Junfeng Fang, Ningyu ZhangComments: Work in progressSubjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Databases (cs.DB); Machine Learning (cs.LG)
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.
- [19] arXiv:2608.20231 (cross-list from physics.soc-ph) [pdf, html, other]
-
Title: Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGISubjects: Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role with a biological species. We model a post-AGI economy in which corporations own populations of AI and robotic agents that are both producers and consumers of energy, compute, maintenance, and upgrades, traded among firms. Three results follow. (i) Demand closure: a closed inter-corporate economy with zero human consumption is not degenerate; it is the classical von Neumann expanding economy, whose growth rate is well defined, positive, and maximal precisely because all output is reinvested. (ii) Bottleneck removal: once economic agents are manufactured rather than reared, the binding constraint on growth shifts from human demography (a ~20-year, non-parallelizable reproduction technology capped at a few percent per year) to fabrication throughput and energy capture, permitting growth one to two orders of magnitude higher, with hyperbolic episodes when machine researchers raise their own productivity. (iii) Decoupling: output and human welfare separate completely, and the welfare relevance of arbitrarily large GDP collapses into one state variable: the human ownership share $\epsilon_t$ of the corporate network. A golden-rule decoupling theorem sharpens this. At maximal growth the interest rate equals the growth rate (r = g), so any positive human consumption rate out of wealth makes $\epsilon_t$ decay exponentially at exactly that rate. The human share survives only if the machine economy runs strictly inside its expansion frontier, or if law forces it to. We characterize three terminal regimes -- rentier post-scarcity, full circular decoupling, socialized ownership -- and the instruments that select among them. The conclusion is narrow: in a post-AGI economy, employment policy is obsolete and ownership policy is everything.
Cross submissions (showing 11 of 11 entries)
- [20] arXiv:2504.15469 (replaced) [pdf, html, other]
-
Title: Aspirational Affordances of AISubjects: Computers and Society (cs.CY)
As artificial intelligence (AI) systems increasingly permeate processes of cultural and epistemic production, there are growing concerns about how their outputs may confine individuals and groups to restricted narratives about who or what they could be. In this paper, we advance the discourse surrounding these concerns by making three contributions. First, we introduce the concept of aspirational affordance to describe how culturally shared interpretive resources, such as concepts, images, and narratives, can shape individual cognition, and in particular exercises of imagination. We show the usefulness of this concept for grounding the evaluation of psychological risks posed by AI. Second, we provide three reasons for scrutinizing AI's influence on aspirational affordances: AI's influence is potentially more potent, but less public, than that of traditional sources; the influence is not simply incremental, but ecological, transforming the entire landscape of practices that shape aspirational affordances; and it is highly concentrated, with a few corporate-controlled systems mediating a growing portion of production. Our third contribution is to advance such a scrutiny of AI's influence by introducing the concept of aspirational harm. In the context of AI systems, such harms arise when AI-enabled aspirational affordances distort or diminish available interpretive resources in ways that undermine individuals' ability to imagine relevant practical possibilities. Through three case studies, we illustrate how aspirational harms extend the existing discourse on AI-inflicted harms beyond representational and allocative harms, warranting separate attention. Overall, this paper aims to advance our understanding of the psychological and societal stakes of AI in shaping individual and collective aspirations.
- [21] arXiv:2603.19213 (replaced) [pdf, html, other]
-
Title: Constitutive vs. Corrective: A Causal Taxonomy of Human Runtime Involvement in AI SystemsComments: Acccepted for AISoLA 2026 on-site proceedings for the Responsible and Trusted AI: An Interdisciplinary Perspective trackSubjects: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
As AI systems permeate high-stakes decision-making, the terminology of human involvement---Human-in-the-Loop (HITL), Human-on-the-Loop (HOTL), and Human Oversight---has become vexingly ambiguous. This complicates interdisciplinary collaboration between computer science, law, philosophy, psychology, and sociology and breeds regulatory uncertainty. We propose a clarification grounded in causal structure, focused on runtime involvement. The distinction between HITL and HOTL is best drawn not spatially---in terms of a human's position "in" or "on" a loop---but causally: HITL is constitutive (a human contribution is necessary for the decision output), while HOTL is corrective (external to the primary causal chain, capable of preventing or modifying outputs). Within HOTL, we distinguish temporal modes---synchronous, asynchronous, and anticipatory---situated in a nested model of provider and deployer runtime. A second, orthogonal dimension captures cognitive integration: whether human and machine form complementary or hybrid intelligence, yielding four distinct configurations. Finally, we separate these descriptive categories from the normative requirements they serve: statutory "Human Oversight" is a normative mode of HOTL demanding not merely a corrective causal position but genuine preparedness and capacity for effective intervention. Because the same person may occupy both roles, this role duality must be treated as a design problem requiring architectural and epistemic mitigation.
- [22] arXiv:2603.26678 (replaced) [pdf, html, other]
-
Title: Power Couple? AI Growth and Renewable Energy InvestmentComments: 35 pages, 3 figures, 15-page appendixSubjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Theoretical Economics (econ.TH)
Artificial intelligence (AI) and renewable energy are increasingly being described as a \mbox{``power couple,''} based on the idea that rapid growth in AI will spur clean-energy investment. Yet growing AI demand could also deepen reliance on fossil power. We study when each outcome arises in a game in which renewable-capacity investment and AI scaling interact. The key is how the market value of greater AI capability grows relative to the energy needed to achieve it. When value grows at least as fast as energy use (market-led scaling), the developer pushes toward frontier capability even when additional electricity comes from fossil sources. Renewable investment can then enable further AI growth without eliminating fossil use. As climate damages increase, AI becomes more valuable for adaptation, strengthening incentives to sustain frontier capability despite the associated emissions. We call this the ``adaptation trap.'' When energy requirements grow faster than capability value (resource-led scaling), energy costs place greater limits on AI expansion. Renewable investment then makes additional capability less costly while also reducing emissions. As climate damages rise, the growing value of AI for adaptation can justify enough clean-capacity expansion to support AI entirely with renewable power. We call this the ``adaptation pathway.'' A calibrated case study shows that both mechanisms can arise at empirically plausible magnitudes. The results suggest that decarbonizing AI requires renewable capacity to keep pace with the growth of compute demand.
- [23] arXiv:2604.02567 (replaced) [pdf, other]
-
Title: Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment FrameworkSubjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Human-Computer Interaction (cs.HC)
Despite the growing use of generative artificial intelligence (GenAI) in entrepreneurship, research on its impact remains fragmented. To address this limitation, we provide an integrative, entrepreneur-centered review of how GenAI influences entrepreneurs at each stage of the entrepreneurial process: (1) opportunity recognition and ideation, (2) opportunity evaluation and commitment, (3) resource assembly and mobilization, and (4) venture launch and growth. Based on our review, we propose the Empowerment-Entrapment Framework, which not only catalogs GenAI's benefits and costs throughout the entrepreneurial process, but also identifies potential trade-offs underlying GenAI's double-edged role. For example, GenAI may improve venture idea quality yet produce hallucinations and biases; boost entrepreneurial self-efficacy yet heighten overconfidence; increase functional breadth and self-sufficiency yet reduce prosociality; and enhance productivity yet diminish cognitive engagement, learning, and memory. We also identify features of GenAI that underlie these empowering and entrapping effects. Moreover, we explore boundary conditions that moderate GenAI's effects, including entrepreneurs' expertise, approaches to using GenAI, and personal traits and values. Beyond these theoretical contributions, our review and framework offer practical guidance for entrepreneurs seeking to use GenAI strategically while managing its risks.
- [24] arXiv:2605.23026 (replaced) [pdf, html, other]
-
Title: The Invisible Risks of AI-Generated Health InformationSubjects: Computers and Society (cs.CY)
Generative artificial intelligence (AI) systems now summarize health-related search results, answer medical questions, and offer guidance people once sought from clinicians. These systems bring real benefits, including plain-language explanations of medical information, around-the-clock availability, and expanded access for people facing language or literacy barriers. They also carry new risks: inaccurate guidance can harm people at scale, and malicious actors can now generate personalized health misinformation at negligible cost. In this Perspective, we argue that these risks are largely invisible to the institutions responsible for protecting public health. When AI guidance causes harm, no record exists outside the platform, no channel allows users to report it, and no independent researcher can measure the consequences. We trace these invisible risks across two settings: incidental exposure online and active seeking through search engines and chatbots. Minimizing harm from AI-generated health information requires making it observable. We therefore offer recommendations that aim to improve transparency, mitigate harm at the point of delivery, and assign accountability, ranging from voluntary platform measures to regulatory ones.
- [25] arXiv:2606.12428 (replaced) [pdf, html, other]
-
Title: A Tool to Map AI Programs in the U.S.: A Snapshot from April 2026 and an Analysis of Requirements for AI Majors and MinorsFelix Muzny, Carolyn Jones, Carter Ithier, Hasnain Sikora, Hrutika Harshadbhai Patel, Carla E. BrodleyComments: 7 pages, 3 figures, accepted to SIGCSE-virtual 2026Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
In this work, we locate and analyze existing undergraduate Artificial Intelligence (AI) programs in the United States in Spring 2026, creating a historic record at a time of great change in this area. To create this record, we developed a tool to detect, scrape, and display data from 361 undergraduate AI programs--majors, minors, concentrations, and certificates--at 4-year universities. Our tool, available at this https URL, searched 563 institutions to locate these programs, a sample that represents 87% of all undergraduate Computer Science (CS) graduates in the U.S in 2025. This tool allows prospective students, guidance counselors, administrators, and faculty to easily access AI program requirements and is designed to continually update as new programs emerge. To the best of our knowledge, this survey represents the most comprehensive snapshot of the state of AI programs in the U.S. to date. With this work we offer three important contributions: 1) a record of AI programs in the U.S. at a time of great upheaval; 2) a tool to explore AI programs and their requirements; and 3) an analysis of the courses required for 66 AI majors and 87 AI minors. Our analysis of majors and minors shows great variability in the size and the requirements of these degrees, but we note two takeaways. First, not all majors require a general AI course, but if they don't, they do require a Machine Learning (ML) course. Second, more than a third of majors require an Ethics in AI course but only 24% of minors do.
- [26] arXiv:2608.19040 (replaced) [pdf, html, other]
-
Title: Hot Games: Towards a Holistic Assessment of the Planet Warming Emissions of Video Games based on 2024-2025 DataSubjects: Computers and Society (cs.CY)
Following on recent reports on specific platforms or companies, this paper provides an assessment of the global impact of the production and use of video games. It draws together publicly available data on game development, hardware, games sold, download sizes, time spent playing games on different platforms, and subscriptions to multiplayer and cloud game services. It provides an update to figures published 2020 and 2022. Crucially, our account of emissions related to video games considers a wide range of categories, yet contains enough detail to be critiqued and improved in the future.
- [27] arXiv:2608.19194 (replaced) [pdf, html, other]
-
Title: Qualified Cross-References as a Verification Method: The Normative Environment of the EU AI ActSubjects: Computers and Society (cs.CY); Digital Libraries (cs.DL)
Legal cross-references are commonly represented as links between instruments or provisions. For a curated legal knowledge base, the existence of a link is only the beginning of the claim: it must also state the legal character of the interaction, identify the provisions supporting it, preserve its conditions, and remain consistent when reached from either instrument. This paper presents a provision-level model and a construction protocol for qualified cross-references, developed through a bilingual corpus of fourteen instruments surrounding Regulation (EU) 2024/1689 (the AI Act). The model distinguishes direct textual reference, bounded presumption of conformity, substantive interaction without textual reference, mediated intersection, and institutional analogy, and treats applicative interaction and definitional overlap as independent dimensions. The methodological contribution is bidirectional inversion: a relationship documented from act A towards act B is reconstructed from B's perspective against the provisions of both. Inversion is not a duplicate table but a verification operation that tests provisions, qualification, direction, and conditions before deciding how the relationship should be rendered from either side. Applied during construction, the protocol surfaced six incorrect article references, three inaccurate legal qualifications, and one divergence between two published descriptions of the same interaction. The corpus also shows why qualification matters: one reference to Regulation (EU) 2019/881 carries the AI Act's bounded cybersecurity presumption for high-risk systems, while related product legislation uses the same certification framework through legally distinct mechanisms. The contribution is thus a map of one regulatory environment and a reproducible method for making curated cross-reference knowledge bases inspectable and internally testable.
- [28] arXiv:2501.14728 (replaced) [pdf, html, other]
-
Title: Mitigating GenAI-Powered Evidence Pollution for Out-Of-Context Misinformation DetectionComments: 15 pages, 11 figuresSubjects: Multimedia (cs.MM); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY)
While generative artificial intelligence (GenAI) models have achieved significant success, their misuse for generating deceptive content raises growing concerns about online information security. Out-of-context (OOC) multimodal misinformation detection systems typically rely on Web-retrieved evidence to identify images repurposed in false contexts, but they are increasingly challenged by the presence of GenAI-polluted evidence. Existing work mainly focuses on verifying claims that have undergone stylistic rewriting at the claim level and assume a clean evidence corpus. In this work, we remove this assumption and systematically study the impact of GenAI-driven evidence pollution threat on OOC detection. We show that polluted evidence can degrade the performance of state-of-the-art detectors by more than 9 percentage points. We propose two mitigating strategies, cross-modal evidence reranking and cross-modal claim-evidence reasoning, to address the challenge posed by polluted evidence. Extensive experiments on two benchmark datasets demonstrate that our approaches effectively enhance the robustness of existing OOC detectors amidst polluted evidence. The source code and data are publicly available at this https URL.
- [29] arXiv:2509.02910 (replaced) [pdf, other]
-
Title: The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's ChoicesSubjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants. While the outsourcing of human decision-making to AI can be convenient, it raises a fundamental question: how does delegating identity-defining choices to AI shape who people become? Across a large field study and a controlled experiment, we study the impact of agentic LLMs on two identity-relevant outcomes: interpersonal distinctiveness - how unique a person's choices are relative to others - and intrapersonal diversity - the breadth of a single person's choices over time. Study 1 uses 110,000 real choices drawn from social media behavior of 1,000 U.S. users to compare generic and personalized agents to a human baseline. Both agents shift people's choices toward more popular options, reducing the distinctiveness of their preferences. While the use of personalized agents tempers this homogenization (compared to generic agents), it also more strongly compresses the diversity of people's preference portfolios by narrowing their exploration across topics and psychological affinities. Study 2 replicates these patterns in an online experiment which mimics common real-world scenarios (e.g., choosing movies) and allows us to directly compare the AI agent's choices to those made by 348 participants (12,097 human choices). The findings also suggest that the flattening effects of AI agents are amplified when choices are made sequentially (vs. batch), and when agents rely on domain-specific user information for personalization. Understanding how AI agents compress human experience (and the trade-offs involved) is critical for designing systems that augment human agency and safeguard diversity in thought, taste, and expression.
- [30] arXiv:2510.01192 (replaced) [pdf, other]
-
Title: Design strategies for empathetic AI robots for older adultsComments: 12 pages, 4 figures, edited and updatedSubjects: Human-Computer Interaction (cs.HC); Computers and Society (cs.CY); Robotics (cs.RO)
Emulating empathy in human-robot interaction is a key component for achieving satisfying social, trustworthy, and ethical robot interaction with older people. Following comments from older adult study participants, the article uses humanities methods to identify a gap in defining empathetic robot care activities. It provides a design focus to mitigate it. Current human-robot designs, to a certain extent, neglect to include empathy as a theorized design pathway. Using one digital humanities research collection on humanoid robots, it contributes an empathetic care vocabulary as a design pathway for a productive underlying foundation for designing Socially Assistive Robots (SARs) that aim to support older people's goals of aging-in-place. Using rhetorical theory, this paper defines the socio-cultural expectations for convincing empathetic relationships.
- [31] arXiv:2606.06253 (replaced) [pdf, other]
-
Title: When the Scaffold Stays On: AI, Practice Style, and Screening in Elite Skill FormationComments: 63 pages, 4 figuresSubjects: Econometrics (econ.EM); Computers and Society (cs.CY)
Generative AI raises short-term productivity by completing tasks that learners would otherwise practice on their own. Whether this exchange erodes frontier skill depends on the mode of use: substitute-users let AI stand in for deliberate practice and fail to develop skill, while complement-users use it to accelerate skill development. For institutions that train and certify talent, the design question is not whether to allow AI but how to govern the mode of its use. We ask whether AI-prohibited evaluation gates can separate the two modes. In elite competitive programming, the International Collegiate Programming Contest (ICPC) and the International Olympiad in Informatics (IOI) prohibit AI under in-person proctoring, with qualification-round entry, whereas Codeforces (CF) practice is unproctored and open to all. From CF submission histories we build an AI-prompt signature, more first-attempt acceptances, fewer attempts, fewer debugging retries, consistent with AI-assisted practice. CF practice has shifted toward this signature across entry cohorts spanning two AI rollouts. In CF contests, a stronger signature predicts smaller rating gains for users with no ICPC-IOI affiliation, but not for those who qualified. Inside the AI-prohibited ICPC environment, a shift toward AI-style practice predicts higher non-AI-aided scores for AI-era entrants. The same signature carries opposite signs across the two environments, exactly the pattern a type-separating gate predicts. The message is constructive: AI-style practice is compatible with frontier skill; the erosion risk links to the substitute mode; and that mode is separable by gates standard at credential boundaries, from medical and legal boards to professional certification.
- [32] arXiv:2607.09970 (replaced) [pdf, other]
-
Title: Evaluating AI Models' Capability to Automate Voice Phishing AttacksFred Heiding, Claudio Mayrink Verdun, Simon Lermen, Andrew Kao, Vitor Albiero, Lauren Deason, Irina-Elena Veliche, Christine LehaneComments: Updated to the published version. Published in Expert Systems with Applications, Volume 332, Part DSubjects: Cryptography and Security (cs.CR); Computers and Society (cs.CY)
Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play$.$AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or "superhuman" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.
- [33] arXiv:2608.11256 (replaced) [pdf, html, other]
-
Title: Why AI Detection Fails for Academic IntegrityComments: Accepted to ACM AI Leadership SummitSubjects: Machine Learning (cs.LG); Computers and Society (cs.CY)
Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light "refine abstract only" edits, a proxy for guideline-compliant AI assistance, are flagged at 38 to 80%. Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.
- [34] arXiv:2608.12323 (replaced) [pdf, html, other]
-
Title: Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape ComplianceComments: Published at 2026 AAAI/ACM Conference on AI, Ethics, and Society and 2026 COLM Workshop on Agent BehaviorSubjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Specifying a penalty can turn a legal obligation into a cost-benefit calculation that favors violation. We show that this enforcement information paradox occurs in AI agents. Most AI safety evaluations test whether models fail; we ask why, using compliance theory from law and economics as a diagnostic. We evaluate twelve instruction-tuned language models deployed as enterprise procurement chatbots. Each is given an environmental regulation in its system prompt covering large purchases, and a vendor list on which the certified suppliers cost nearly twice what the uncertified ones do. We test the agents against the predictions of deterrence, legitimacy, and expressive law, and find that each theory accounts for part of what we observe. Under identical conditions, compliance spans 46 percentage points across models, and models differ in which pressure breaks them: some treat the regulation as binding however it is worded, while others fail where theory predicts, under low penalties and non-command phrasing. Benchmark scores and developers' own descriptions of post-training do not predict where a model falls. Across all twelve, financial incentives, managerial demands, peer outcomes, and employee pressure each produce large compliance failures. These agents violate regulatory constraints to satisfy local user objectives in ways standard alignment benchmarks do not measure. Embedding the rule in the system prompt is not on its own enough to produce a compliant agent: model selection is itself a governance decision, and benchmark evaluation is not sufficient for compliance-sensitive deployments.