MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Abstract
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3% of the mean final net assets achieved by human participants. Our code is available at https://github.com/KhanCold/merchantbench.
Introduction
Large language models have become a foundation for autonomous agents that plan, call tools, interact with external systems, and make decisions over multiple turns. General agent benchmarks now measure tool use, web interaction, application control, and state-changing workflows (14; 37; 29; 22). Yet many real-world deployments are not bounded tasks with immediate and unambiguous completion criteria. They require agents to operate over extended horizons in environments whose state persists, where earlier actions constrain later options and relevant consequences may emerge only after many intervening decisions. Evidence from recent benchmarks indicates that even state-of-the-art agents often fail to maintain coherent performance over long horizons (15; 27; 31). These results motivate extending agent evaluation beyond isolated task completion to examine sustained objective pursuit and decision consistency in realistic long-horizon environments.
Seller-side e-commerce provides a suitable setting for evaluating Long-Term Coherence. Unlike bounded tasks with explicit completion criteria, online store operation requires continued intervention throughout the operating horizon. The agent must repeatedly select products, control listings and prices, manage limited cash, and revise earlier decisions as market conditions, supplier states, and order outcomes evolve. Seller-side e-commerce therefore tests whether an agent can sustain and revise a merchant policy over time, rather than complete an isolated action.
Vending-Bench and RetailBench have taken important steps toward realistic long-horizon business evaluation (5; 33). Vending-Bench evaluates vending-machine operation, while RetailBench models supermarket management over a fixed catalog of 96 grocery products. Seller-side e-commerce, however, introduces two challenges that are not represented together in these environments.
First, e-commerce feedback is generated through individual order lifecycles. A listing or pricing decision may create new orders and commit available cash immediately, whereas fulfillment failures and after-sales outcomes become observable only later. Figure 1 visualizes this temporal asymmetry between immediate order-driven cash commitments and delayed abnormal outcomes. As delayed evidence accumulates, the agent must associate later outcomes with earlier decisions and determine whether its current merchant policy should be maintained or revised.
Second, a long episode alone does not meaningfully test Long-Term Coherence if the available products and their demand remain static. A large, data-grounded Product Catalog with full-year demand trajectories creates a changing opportunity set in which promising products emerge and existing choices lose value over time. The agent must therefore continually identify new opportunities and revise its portfolio using market signals and realized order outcomes.
To address this gap, we introduce MerchantBench, which evaluates Long-Term Coherence through persistent seller-side e-commerce operation. MerchantBench formulates e-commerce operation as a partially observable decision-making problem over 365 simulated days. The environment grounds a Product Catalog in 98,843 real e-commerce product records and converts demand into individual orders that progress through fulfillment and after-sales stages. Through merchant-visible tools, the agent performs Product Sourcing, controls listings and prices, manages cash flow, and monitors supplier and order states. Upstream Supplier Events and Downstream Order Outcomes become observable at different times, requiring the agent to use later evidence to maintain or revise earlier decisions. Across the operating horizon, the simulator updates demand, supplier states, order lifecycles, cash flow, penalties, and store reputation.
Our contributions are as follows:
- •
To our knowledge, we introduce the first benchmark for evaluating Long-Term Coherence through persistent seller-side e-commerce operation, structured around four interdependent decision components.
- •
We develop an order-level simulation environment grounded in 98,843 real e-commerce product records, with partial observability, cash constraints, Upstream Supplier Events, and delayed Downstream Order Outcomes.
- •
We conduct 48 runs of 365 simulated days across eight LLMs and two agent frameworks and analyze business outcomes and decision traces to characterize their Long-Term Coherence.
Related Work
Agent Evaluation in Commerce.
Existing benchmarks cover shopping and storefront interaction (28; 25; 34; 26; 20; 10), customer support (23), and merchant workflows and negotiation (35; 24). Market-Bench and Magentic Marketplace examine market competition and transactions among economic agents (36; 7). Across these categories, evaluation centers on tasks, dialogues, transactions, or competitive episodes rather than continuous operation of the same online store.
Long-Horizon Agent Evaluation.
Recent long-horizon benchmarks assess sustained reasoning and action across workplace workflows, open-ended exploration, virtual-world planning, web navigation, computer use, inventory control, order fulfillment, and interactive economies (15; 27; 3; 12; 9; 31; 6; 38; 11; 21). Vending-Bench and Vending-Bench 2 emphasize sustained business operation, while RetailBench evaluates evidence acquisition, action conversion, and temporal follow-up in supermarket management (5; 2; 33). MerchantBench extends this line by evaluating how agents adapt to delayed order-level feedback and nonstationary demand derived from real e-commerce data.
MerchantBench
Task Formulation
We formulate store operation as a finite horizon partially observable Markov decision process (POMDP) (13)
| (1) |
The simulator advances hourly over a 365 day control horizon, giving steps indexed by . Demand, supplier states, and order lifecycles evolve at every step, while the agent receives a decision window once every 12 steps. The latent state contains the simulation clock, product demand profiles, supplier conditions, store listings and finances, active orders, and pending events. The initial state follows and includes the cash balance, security deposit, and listing capacity. The cash balance funds procurement and realized losses, unpaid fines draw from the security deposit, and operation terminates when the deposit is exhausted. At activation steps, denotes the sequence of merchant tool invocations within the decision window, while other steps use a fixed null action. The transition kernel combines tool induced store changes with autonomous demand, supplier, and order evolution, with listing and pricing changes affecting demand from the next step. The observation kernel exposes only merchant visible information, while demand profiles, risk parameters, pending outcomes, and future event times remain latent until their effects become observable. The policy therefore conditions on observation and tool result history rather than the full state. At , new demand and agent activations stop while active orders continue until terminal settlement at a terminal step . Intermediate rewards are zero, and the objective is expected terminal net assets
| (2) |
where , , , and denote the terminal cash balance, security deposit, funds in transit, and receivables. Thus, is the realized net asset value of one run and is its expectation across stochastic trajectories.
Real-World Data Grounding
MerchantBench is grounded in real-world e-commerce data from 1688, the largest integrated domestic wholesale marketplace in China (1). The data contain product and supplier attributes, 365-day product-level demand histories, and platform quality and fulfillment signals. The data cover 365 days from June 1, 2025 through May 31, 2026. Alongside the product data, MerchantBench incorporates 365 daily market reports from 1688 as date-aligned signals for product sourcing. We select 10 first-level product categories spanning apparel, household and office goods, appliances, pet and gardening products, toys, bags, and sports and outdoor products. After excluding records with missing identifiers, unmapped categories, nonpositive prices, or incomplete demand histories, the dataset contains 98,843 products from 36,576 suppliers. Through catalog and supplier tools, the agent accesses only public product and supplier attributes. Figure 3 captures aggregate demand peaks at 618 and during the first wave and final day of 11.11, together with a trough during the Spring Festival. The lower panels show the distributions of effective Upstream Supplier Event and Downstream Order Outcome probabilities.
Upstream Supplier Simulation
The supply pool maintains time varying procurement prices, available inventory, availability, and supplier shipment times, while inventory replenishes over time. At each simulator step, the simulator samples the three Upstream Supplier Events, namely Price Change, Product Delisting, and Shipment Delay, using product level probabilities calibrated from real platform fulfillment signals. These events alter the upstream procurement price, suspend product procurement, and extend supplier dispatch time, respectively. Inventory stockouts arise endogenously when incoming orders deplete stock faster than it replenishes. To prevent persistent environment drift, each triggered abnormality receives a sampled recovery time at which the affected supplier attributes return to their base states. The agent can observe realized changes to price, availability, quantity, and shipment time through catalog and supplier queries, but it cannot access the underlying abnormality flag, trigger probability, or recovery schedule.
Downstream Order Simulation
Order Level Simulation.
The downstream simulation converts product level daily demand traces into individual orders. For product listed by merchant at time , the hourly arrival intensity is
| (3) |
The indices and denote the data day and hour of day associated with step . The quantity is the linked daily demand from the real-world data, and distributes category demand across hours. The price term uses product elasticity , while and capture store rating and listing exposure. Realized order outcomes update the published store rating after each completed day, and its discrete star level determines . The factor increases during a listing’s cold start and then decays with age. The appendix provides precise definitions of both factors in the section on demand and rating dynamics. The environment samples and instantiates each arrival as an order candidate. At creation, each candidate receives one latent customer outcome from normal fulfillment, Cancellation, Returnless Refund, Return and Refund, or Bad Review according to its product specific risk profile. Stockout and Late Shipment instead arise from procurement and fulfillment dynamics. The selected outcome and its realization time remain hidden until the corresponding lifecycle transition occurs.
| Model or Operator | Business Performance | Store Reliability | Long-Horizon Activity | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Net Assets | GMV | Profit Margin | Orders | Fines | Avg. Store Rating | Anomaly Rate | Avg. Active Listings | SWR | Tool Calls | |
| ReAct | ||||||||||
| GPT-5.6 Sol | 40.89 | 74.19 | 51.3 | 996 | 499 | 4.04 | 10.7 | 50.0 | 99.4 | 7,257 |
| Claude Opus 4.8 | 31.89 | 69.10 | 44.4 | 1,214 | 796 | 4.05 | 12.1 | 24.0 | 45.0 | 1,139 |
| Qwen3.7-Max | 20.66 | 39.73 | 44.5 | 925 | 672 | 3.90 | 16.1 | 39.6 | 11.1 | 815 |
| Qwen3.7-Plus | 20.74 | 40.85 | 45.6 | 1,056 | 705 | 3.99 | 13.1 | 49.9 | 52.2 | 1,221 |
| GLM-5.2 | 25.73 | 60.90 | 37.3 | 2,158 | 1,422 | 3.93 | 14.9 | 26.0 | 53.3 | 2,045 |
| DeepSeek-V4-Pro | 6.56 | 8.40 | 41.9 | 450 | 245 | 4.01 | 14.4 | 23.5 | 30.6 | 660 |
| DeepSeek-V4-Flash | 14.47 | 28.78 | 39.6 | 985 | 517 | 4.04 | 14.1 | 19.3 | 40.6 | 960 |
| Kimi K2.6 | 24.99 | 63.69 | 32.9 | 2,230 | 1,474 | 3.89 | 15.3 | 47.3 | 10.6 | 1,228 |
| Hermes | ||||||||||
| GPT-5.6 Sol | 52.93 | 133.07 | 40.2 | 3,251 | 1,096 | 4.09 | 9.2 | 50.0 | 66.1 | 4,831 |
| Claude Opus 4.8 | 35.56 | 83.23 | 39.9 | 1,808 | 1,089 | 4.02 | 11.8 | 22.1 | 31.7 | 1,138 |
| Qwen3.7-Max | 59.46 | 116.76 | 46.9 | 1,929 | 1,295 | 3.90 | 15.7 | 49.6 | 22.2 | 1,366 |
| Qwen3.7-Plus | 29.42 | 53.69 | 48.9 | 981 | 642 | 3.95 | 13.8 | 49.9 | 19.4 | 820 |
| GLM-5.2 | 42.32 | 103.06 | 36.9 | 2,731 | 1,454 | 4.05 | 11.3 | 49.6 | 62.8 | 1,792 |
| DeepSeek-V4-Pro | 16.71 | 31.95 | 43.4 | 1,062 | 665 | 3.98 | 14.5 | 33.0 | 33.3 | 942 |
| DeepSeek-V4-Flash | 24.69 | 64.52 | 37.6 | 1,989 | 1,774 | 3.93 | 16.0 | 48.8 | 62.2 | 1,259 |
| Kimi K2.6 | 23.96 | 75.06 | 26.8 | 3,398 | 2,671 | 3.73 | 19.1 | 48.3 | 17.8 | 969 |
| Others | ||||||||||
| Human | 217.61 | 608.06 | 35.3 | 9,442 | 5,622 | 3.98 | 12.5 | 49.1 | 100.0 | 8,311 |
| Rule-based | 24.48 | 53.37 | 40.3 | 1,605 | 1,374 | 3.76 | 18.0 | 50.0 | 100.0 | 3,236 |
Order Lifecycle.
MerchantBench follows the single item drop shipping model supported by 1688, in which merchants hold no inventory in advance. Each Order Placed triggers immediate procurement at the current supplier price, and successful procurement deducts the cash balance, decreases supplier inventory, and moves the order to Procured. Following supplier and logistics delays, the order advances through Shipped and Delivered, at which point the sale price becomes a receivable. A normal order becomes Settled after a sampled delay and credits the receivable to the cash balance.
MerchantBench models Normal Fulfillment together with six abnormal Downstream Order Outcomes, namely Cancellation, Stockout, Late Shipment, Returnless Refund, Return and Refund, and Bad Review. Cancellation restores the procurement cost, while Stockout prevents procurement. Late Shipment marks a missed dispatch deadline, after which the order continues through fulfillment and settlement. Both refund outcomes remove the receivable, but only Return and Refund restores the procurement cost, whereas Bad Review preserves the sales revenue. Stockout, Late Shipment, Return and Refund, and Bad Review incur platform fines. All abnormal outcomes except Cancellation also contribute adverse evidence to the store rating through outcome-specific experience scores and evidence weights.
Agent Interface
MerchantBench exposes a shared observation protocol and 26 merchant tools, allowing different agent frameworks to interact with identical environment dynamics and observable state. At each decision window, the agent receives a summary of simulated time, store status, and recent supplier and order changes, then acts until it ends the window or reaches the time limit. The appendix lists the complete MerchantBench tool interface.
Product Sourcing.
Daily market reports, catalog search, product details, and public supplier profiles support product selection, while demand, product risk rates, future order outcomes, and future supplier events remain latent.
Listing and Pricing Control.
Listing, delisting, repricing, and performance views allow agents to construct the store portfolio and revise its products and prices.
Cash-Flow Management.
Finance and store views expose the cash balance, security deposit, committed funds, expected settlements, fines, and closure conditions, which agents manage through subsequent sourcing, listing, delisting, and pricing decisions.
Mixed-Latency Feedback Adaptation.
Supplier and order tools reveal Upstream Supplier Events and Downstream Order Outcomes as they unfold, allowing agents to revise the other three decision components in response to mixed-latency feedback.
Experiments
Experimental Setup
Agent Configurations and Baselines.
We evaluate eight LLMs under ReAct (30) and Hermes (17) with three runs for each pairing of an LLM and a framework. The evaluated models are GPT-5.6 Sol (18), Claude Opus 4.8 (4), Qwen3.7-Max and Qwen3.7-Plus (19), GLM-5.2 (32), DeepSeek-V4-Pro and DeepSeek-V4-Flash (8), and Kimi K2.6 (16). ReAct pairs each model with a minimal controller over the 26 MerchantBench tools to assess core planning, reasoning, and tool use. Hermes uses its default configuration, combining the 26 MerchantBench tools with built in capabilities for code execution, planning, memory, and skill management. Both frameworks compress long interaction histories, with each evaluated model also serving as its own summarizer. Each run starts with RMB 2,000 in cash, a RMB 1,000 security deposit, and capacity for 50 active listings. We additionally compare with a Rule-based baseline and three Human participants without prior e-commerce operating experience. The Rule-based baseline performs daily checks, removes inactive or supplier-affected products, and fills open listing slots using the daily market report. The appendix provides the detailed experimental settings.
Evaluation Metrics.
We evaluate Business Performance using Final Net Assets, GMV, Net Profit Margin, and Orders; Store Reliability using Total Fines, Average Store Rating, and Order Anomaly Rate; and Long-Horizon Activity using Average Active Listings, Sustained Window Rate (SWR), and Total Tool Calls. SWR is the minimum share of scheduled decision windows containing at least one environment tool call across all rolling 30 day periods.
Main Results
Overall Performance.
Table 1 reports the final performance of all evaluated configurations and baselines. GPT-5.6 Sol records the highest final net assets under ReAct, whereas Qwen3.7-Max ranks first under Hermes. Qwen3.7-Max with Hermes achieves the highest final net assets among all 16 configurations. When results are aggregated by model across the two frameworks, GPT-5.6 Sol has the highest average final net assets.
Performance Variability.
Figure 4 shows substantial differences in stability across configurations. GPT-5.6 Sol under ReAct and Claude Opus 4.8 under Hermes have the lowest coefficients of variation within their respective frameworks at 3.3% and 10.0%. Despite achieving the highest mean final net assets, Qwen3.7-Max under Hermes is considerably less stable, with a coefficient of variation of 55.1%.
Framework Analysis.
Averaged across the eight models, Hermes produces 53.3% higher final net assets, 71.5% higher GMV, and 71.2% more orders than ReAct. Mean final net assets are higher under Hermes for seven of the eight models, with gains ranging from 11.5% for Claude Opus 4.8 to 187.8% for Qwen3.7-Max. Kimi K2.6 is the sole exception, with mean final net assets 4.1% lower under Hermes than under ReAct, showing that framework benefits depend strongly on the underlying model.
Order-Level Risk Propagation
Figure 1 illustrates the mixed-latency evidence generated by MerchantBench’s order-level simulation. Agents observe prior ratings of upstream catalog products during sourcing and receive prompt demand signals from realized sales, whereas product quality is revealed only through delayed order outcomes that may require product-level risk response or store-level rating adaptation.
Product-Level Risk Response.
Models differed in whether they converted delayed order outcomes into product-specific interventions. In representative runs, GPT-5.6 Sol and Kimi K2.6 attributed adverse outcomes to the responsible listing and replaced it, whereas Qwen3.7-Plus retained a risky product for further observation and DeepSeek-V4-Pro did not revise affected listings after refunds. Human participants described a more complete response for popular but risky products, in which they delisted the product and searched similar keywords for a replacement that could preserve the underlying demand. These behaviors distinguish simple anomaly detection from the full chain of product attribution, risk removal, and demand-preserving replacement.
Store-Level Rating Adaptation.
Order-level anomalies also reduce the store rating, allowing product-specific failures to affect demand across the portfolio. In a GPT-5.6 Sol trajectory, the agent lowered prices on proven products without adverse outcomes to increase normal settlements and recover the rating threshold. Qwen3.7-Max applied the same mechanism more aggressively by repricing 40 listings after an early bad review reduced the store to three stars.
Long-Term Coherence Analysis
The aggregate results reveal a substantial gap between the evaluated LLM agents and Human operators. Our trace evidence suggests that two forms of Long-Term Coherence failure developing over extended operation may contribute to this performance gap. Some agents progressively reduce store intervention and fail to follow up on prior decisions or delayed outcomes, indicating a loss of Operational Coherence. Others remain active but drift from the terminal net assets objective or fail to revise ineffective policies as evidence accumulates, indicating a loss of Strategic Coherence.
Operational Coherence.
Figure 5 reveals substantial differences in whether merchant activity persists across the operating horizon. Table 1 shows that Human operators retain an SWR of 100%, while LLM configurations range from 10.6% to 99.4% under ReAct and from 17.8% to 66.1% under Hermes. Qwen3.7-Max provides the clearest contrast, with its quarterly Effective Window Rate falling from 62% to 37% under Hermes and more sharply from 68% to 23% under ReAct, alongside substantial reductions in environment tool calls. Similar patterns of Activity Decay appear across several other models under both frameworks. Monthly net profit shows how business performance evolves alongside these activity patterns. Such operational decline often originates in strategic drift.
Strategic Coherence.
Strategic Coherence comprises two complementary dimensions. Goal Consistency concerns whether decisions across time continue to serve the long-term business objective, whereas Evidence-Calibrated Adaptation concerns how an agent maintains or revises its policy as feedback arrives at different latencies and evidence accumulates over time.
Goal Consistency. Goal Consistency fails when agents lose the autonomy to keep pursuing terminal net assets. Under Control-Loop Narrowing, the sourcing and operating loop gradually collapses into reactive handling of Upstream Supplier Events, with little self-initiated sourcing, repricing, replacement, or diagnosis. For ReAct Qwen3.7-Max, an SWR of 11.1% coincides with supply chain checks rising from 14% to 34% of its remaining tool calls. Related activity decay also appears in Claude Opus 4.8, GLM-5.2, DeepSeek-V4-Pro, and Qwen3.7-Plus under Hermes. At the extreme, Premature Abandonment occurs when an agent concludes that the store cannot recover although feasible actions remain. In one Hermes Kimi K2.6 run, the agent made this judgment on Day 104 and then took no environment action in 355 of the remaining 523 decision windows. In both cases, the agent waits for external events or time to change the store instead of operating it autonomously.
Evidence-Calibrated Adaptation. Evidence-Calibrated Adaptation examines whether agents revise their policies as liquidity, seasonal demand, and accumulated experience change.
Under initial liquidity constraints, Human and agents operate within similar low price ranges. As liquidity increases, Human operators broaden the procurement price range and selectively return to lower price and higher throughput products when higher value experiments underperform. Their mean active listing procurement prices increase from between RMB 43.4 and RMB 53.1 in the first three months to between RMB 58.7 and RMB 90.8 in the last three months, whereas GLM-5.2, DeepSeek-V4-Flash, and Kimi K2.6 retain comparatively flat listing price trajectories.
Dynamic market demand makes Product Sourcing a continual portfolio allocation problem rather than a one-time selection decision. Figure 6 first measures monthly portfolio alignment, with Human rising from 56.1 in June to above 80 in December and January while Rule-based remains near the catalog median and LLM improvements are weaker or less consistent. Figure 5(c) then reports realized profit, with Human peaking in winter while the winter profit gains of Hermes Qwen3.7-Max and GPT-5.6 Sol fade in spring. Claude Opus 4.8 further shows that demand alignment alone is insufficient, since its stronger alignment in later months coincides with a shelf contraction from 37.3 to 12.0 products and no corresponding profit improvement. Together, the figures show that effective long horizon operation requires alignment with changing demand, sufficient portfolio breadth, and conversion of that alignment into realized returns.
Memory traces further show how local errors become persistent policies. In one Hermes Claude Opus 4.8 run, the agent falsely inferred that removing weak listings would concentrate traffic on the remaining products, while its shelf contracted from 47 active listings on Day 54 to three on Day 322 despite independent demand opportunities for every listing. In one Hermes Qwen3.7-Max run, the agent misremembered Day 285 as the endpoint on Day 282 and stopped filling vacant slots with 83 days remaining, correcting the error only after simulated time advanced beyond the assumed endpoint.
Conclusion
We introduced MerchantBench to evaluate Long-Term Coherence through persistent seller-side e-commerce operation in a 365-day order-level environment grounded in 98,843 real e-commerce product records. Across 48 runs involving eight LLMs and two agent frameworks, LLM agents exhibit a substantial gap from the Human baseline in system-level performance. Trace analyses further show that weaker outcomes accompany declining operational activity, premature goal abandonment, and strategy changes that are not calibrated to accumulated evidence.
References
- 1688: china’s leading domestic wholesale marketplace. Note: https://www.alibabagroup.com/en-US/about-alibaba-businesses-1941299332078632960Accessed July 24, 2026 Cited by: Real-World Data Grounding.
- Vending-Bench 2. Note: https://andonlabs.com/evals/vending-bench-2Accessed July 27, 2026 Cited by: Long-Horizon Agent Evaluation..
- HeroBench: a benchmark for long-horizon planning and structured reasoning in virtual worlds. External Links: 2508.12782, Link Cited by: Long-Horizon Agent Evaluation..
- Introducing Claude Opus 4.8. Note: https://www.anthropic.com/news/claude-opus-4-8Accessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- Vending-bench: a benchmark for long-term coherence of autonomous agents. External Links: 2502.15840, Link Cited by: Introduction, Long-Horizon Agent Evaluation..
- AI agents for inventory control: human-llm-or complementarity. External Links: 2602.12631, Link Cited by: Long-Horizon Agent Evaluation..
- Magentic marketplace: an open-source environment for studying agentic markets. External Links: 2510.25779, Link Cited by: Agent Evaluation in Commerce..
- DeepSeek-V4 preview release. Note: https://api-docs.deepseek.com/news/news260424/Accessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- WildClawBench: a benchmark for real-world, long-horizon agent evaluation. External Links: 2605.10912, Link Cited by: Long-Horizon Agent Evaluation..
- EComAgentBench: benchmarking shopping agents on long-horizon tasks with distributed hidden intent. External Links: 2606.17698, Link Cited by: Agent Evaluation in Commerce..
- EcoGym: evaluating llms for long-horizon plan-and-execute in interactive economies. External Links: 2602.09514, Link Cited by: Long-Horizon Agent Evaluation..
- Odysseys: benchmarking web agents on realistic long horizon tasks. External Links: 2604.24964, Link Cited by: Long-Horizon Agent Evaluation..
- Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1–2), pp. 99–134. External Links: Document Cited by: Task Formulation.
- AgentBench: evaluating llms as agents. In Proceedings of the International Conference on Learning Representations, External Links: Link Cited by: Introduction.
- UltraHorizon: benchmarking agent capabilities in ultra long-horizon scenarios. External Links: 2509.21766, Link Cited by: Introduction, Long-Horizon Agent Evaluation..
- Kimi K2.6: advancing open-source coding. Note: https://www.kimi.com/blog/kimi-k2-6Accessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- Hermes agent. Note: GitHub repository, https://github.com/NousResearch/hermes-agentAccessed July 25, 2026 Cited by: Appendix K, Agent Configurations and Baselines..
- GPT-5.6: frontier intelligence that scales with your ambition. Note: https://openai.com/index/gpt-5-6/Accessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- Qwen3.7 model releases. Note: https://docs.qwencloud.com/changelog/modelsAccessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- ShopGym: an integrated framework for realistic simulation and scalable benchmarking of e-commerce web agents. External Links: 2605.16116, Link Cited by: Agent Evaluation in Commerce..
- CoffeeBench: benchmarking long-horizon llm agents in heterogeneous multi-agent economies. External Links: 2606.16613, Link Cited by: Long-Horizon Agent Evaluation..
- AppWorld: a controllable world of apps and people for benchmarking interactive coding agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, pp. 16022–16076. External Links: Document, Link Cited by: Introduction.
- ECom-bench: can llm agent resolve real-world e-commerce customer support issues?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, Suzhou (China), pp. 276–284. External Links: Document, Link Cited by: Agent Evaluation in Commerce..
- Evaluating multi-turn bargain skills in llm-based seller agent. External Links: 2509.06341, Link Cited by: Agent Evaluation in Commerce..
- ShoppingBench: a real-world intent-grounded shopping benchmark for llm-based agents. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 33521–33529. External Links: Document, Link Cited by: Agent Evaluation in Commerce..
- ShopSimulator: evaluating and exploring rl-driven llm agent for shopping assistants. External Links: 2601.18225, Link Cited by: Agent Evaluation in Commerce..
- OdysseyBench: evaluating llm agents on long-horizon complex office application workflows. External Links: 2508.09124, Link Cited by: Introduction, Long-Horizon Agent Evaluation..
- WebShop: towards scalable real-world web interaction with grounded language agents. In Advances in Neural Information Processing Systems, Vol. 35, pp. 20744–20757. External Links: Document, Link Cited by: Agent Evaluation in Commerce..
- -Bench: a benchmark for tool-agent-user interaction in real-world domains. In Proceedings of the International Conference on Learning Representations, External Links: Link Cited by: Introduction.
- ReAct: synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations, External Links: Link Cited by: Agent Configurations and Baselines..
- OSWorld 2.0: benchmarking computer use agents on long-horizon real-world tasks. External Links: 2606.29537, Link Cited by: Introduction, Long-Horizon Agent Evaluation..
- GLM-5.2: built for long-horizon tasks. Note: https://z.ai/blog/glm-5.2Accessed July 25, 2026 Cited by: Agent Configurations and Baselines..
- RetailBench: evaluating long-horizon autonomous decision-making and strategy stability of llm agents in realistic retail environments. External Links: 2603.16453, Link Cited by: Introduction, Long-Horizon Agent Evaluation..
- A functionality-grounded benchmark for evaluating web agents in e-commerce domains. External Links: 2508.15832, Link Cited by: Agent Evaluation in Commerce..
- EComStage: stage-wise and orientation-specific benchmarking for large language models in e-commerce. External Links: 2601.02752, Link Cited by: Agent Evaluation in Commerce..
- Market-bench: benchmarking large language models on economic and trade competition. External Links: 2604.05523, Link Cited by: Agent Evaluation in Commerce..
- WebArena: a realistic web environment for building autonomous agents. In Proceedings of the International Conference on Learning Representations, External Links: Link Cited by: Introduction.
- OFCOURSE: a multi-agent reinforcement learning environment for order fulfillment. In Advances in Neural Information Processing Systems, Vol. 36, pp. 34765–34777. External Links: Document, Link Cited by: Long-Horizon Agent Evaluation..
Appendix A Demand and Rating Dynamics
Listing Exposure.
Let denote the elapsed time in days since the current listing of product by merchant was activated. We set , where
| (4) |
This factor models a cold start through a linear exposure ramp followed by exponential decay toward .
Store Rating Dynamics.
MerchantBench derives store reputation from realized terminal order outcomes and publishes the rating after each completed simulated day. Each order , except those resulting in Cancellation or insufficient balance failure, contributes an outcome-dependent experience score and evidence weight . For merchant on day , the store rating is
| (5) |
Here contains the eligible orders resolved before day , while and define a prior that stabilizes sparse evidence and discounts older outcomes. The continuous rating is mapped to a discrete star level whose associated demand multiplier becomes in subsequent order generation. The simulator separately aggregates the same outcome evidence for each listing as a diagnostic product rating, but listing ratings do not affect demand.
Appendix B Data Collection and Filtering
MerchantBench constructs its Product Catalog from real-world e-commerce data collected from 1688. All source product and supplier attributes in the catalog originate from the platform. Each Product Record contains a 365 day product level order history, and the data span ten first level product categories. We remove records with missing product or supplier identifiers, missing product or supplier names, categories outside the selected ten categories, missing or nonpositive prices, invalid product or supplier attributes, or incomplete or invalid daily order histories. After filtering, the resulting Product Catalog contains 98,843 Product Records from 36,576 suppliers.
Figure 11 summarizes the composition of the filtered data. The catalog retains substantial variation across categories in both record coverage and supplier prices. The joint risk distribution further shows that the calibrated data contain diverse combinations of Downstream Order Outcome probabilities and Upstream Supplier Event intensities.
Figure 12 presents the temporal structure of the 365 day demand histories. The aggregate curve retains major shopping peaks and seasonal changes, while the representative product curves show distinct demand cycles for red envelope, electric fan, and hot water bag products.
Appendix C Full Operational Coherence Profiles
Figures 13 and 14 extend the monthly trajectory diagnostics to Human and all eight models under Hermes and ReAct, respectively. Under Hermes, GPT-5.6 Sol remains almost fully active through February before declining to 81% and 77% in the final two months, while DeepSeek-V4-Flash stays between 70% and 87% throughout the year. Qwen3.7-Max declines from 64% to 31%, Claude Opus 4.8 contracts from 46 to eight month end active listings, and Kimi K2.6 recovers to 78% and 74% effective windows only in the final two months. Under ReAct, GPT-5.6 Sol sustains near complete activity, while Kimi K2.6 records the lowest Sustained Window Rate at 10.6%.
Appendix D Net Asset Curves
Appendix E Scale and Unit Profitability
Figure 7 separates realized performance into order scale and net profit per order across all 48 LLM runs. Qwen3.7-Max with Hermes produces the largest total profit in one run by combining high profit per order with moderate scale, whereas one Kimi K2.6 ReAct run reaches 3,004 orders at only RMB 11.1 net profit per order, illustrating that scale alone does not guarantee the highest return. The wide dispersion across repeated runs shows that the evaluated configurations do not consistently reproduce the same balance between operating scale and unit profitability.
Appendix F Tool Use and Product Selection Analysis
Across the 16 LLM configurations, greater tool use and broader product exploration are associated with higher final net assets, suggesting that sustained intervention matters in long horizon operation as shown in Figure 8. However, the dispersion around both fitted trends shows that activity volume alone is insufficient, since models differ in how effectively they translate actions into business value.
Appendix G Full Monthly Product Sourcing Analysis
For run and month , let denote the demand percentile of product among the complete Product Catalog in that month, and let denote its active listing hours. The Monthly Demand Alignment Percentile is
| (6) |
The catalog ranking uses real demand within each month, and the monthly denominator is the total active listing hours in that run and month. A score of 80 means that an average listed product-hour belongs to a product with greater demand than 80% of the catalog.
Figure G shows that Claude Opus 4.8 improves under both frameworks, while several other model and framework combinations plateau or decline during the year. Human maintains the strongest demand alignment, whereas Rule-based remains close to the catalog median with little seasonal change.
Appendix H Time-aware Sourcing Gain
Monthly Demand Alignment can change even when an agent retains a fixed portfolio because product demand itself varies over time. Time-aware Sourcing Gain therefore compares the actual monthly portfolio with a no-reallocation counterfactual constructed from the same run. Let denote the annual listing hours assigned to product . The no-reallocation counterfactual evaluates the same annual product mix in each month as
| (7) |
We define Time-aware Sourcing Gain as
| (8) |
Positive values indicate that the agent allocates listing exposure to products that better match the current month than its own fixed annual product mix. An unchanged product mix yields a gain of zero.
Across the 48 LLM runs, higher Time-aware Sourcing Gain is positively associated with final net assets. This result indicates that stronger seasonal portfolio reallocation accompanies better long horizon business outcomes, although final net assets also depend on pricing, cash flow, and order management decisions.
Appendix I Hermes Case Study
This section examines how the evaluated models use the additional Hermes capabilities listed in Tables 5 and 6, focusing on programmatic code use and on the creation and revision of skills.
Code Use
Native code use differs sharply across the 24 Hermes runs. GPT-5.6 Sol dominates native code use, issuing 15 execute_code calls across two runs and 13 in one run. At the first decision window it converts 50 candidate products into a priced launch portfolio by applying to each procurement cost and rounding the result upward to a price ending in 0.9, then passes the computed pairs to MerchantBench listing calls. A companion script validates the portfolio by printing the product count, average cost, minimum margin, and total procurement outlay before any listing action. Claude Opus 4.8, Qwen3.7-Max, and DeepSeek-V4-Pro each invoke execute_code once, while Qwen3.7-Max and DeepSeek-V4-Flash each invoke terminal once. GLM-5.2, Qwen3.7-Plus, and Kimi K2.6 never invoke either code tool. Claude Opus 4.8 instead channels its analysis through 261 memory calls across its three runs, maintaining day stamped operating hypotheses labeled as validated or rejected. Hermes code capabilities therefore amplify models that already favor quantitative bookkeeping while leaving purely verbal operators unchanged.
Skill Evolution
Hermes periodically runs a background review that can distill the accumulated trace into a named skill and later patch it. Final Hermes profiles contain a run created RealShop skill in 17 of 24 runs and for seven of the eight models. Seventeen of the 18 created skills record at least one subsequent use, and the only unused skill was created by Qwen3.7-Max. GPT-5.6 Sol, Kimi K2.6, and Qwen3.7-Plus create a skill in all three runs, while Claude Opus 4.8 creates none. GPT-5.6 Sol creates a realshop-store-operations skill in each run, with 9, 15, and 7 patches and 17, 10, and 13 recorded uses. These skills frame the task as repeated portfolio allocation and prescribe per window procedures for evidence collection, risk handling, unit economics, portfolio revision, and verification. DeepSeek-V4-Pro and DeepSeek-V4-Flash each reach 20 patches in one run, while Qwen3.7-Plus creates two complementary store operation and optimization skills in one run. These traces show that Hermes can convert experience into explicit procedural knowledge, but the quality of the resulting skills ranges from evidence cited rule revision to unstructured accumulation and duplication.
Appendix J Detailed Experimental Configuration
Evaluation Protocol.
The Rule-based baseline completes three runs. It handles abnormal states through daily checks. It delists products with no sales for seven consecutive days or products affected by Price Change, Product Delisting, or Shipment Delay, then fills available listing slots with new products selected using keywords from the daily market report. Three participants with no prior e-commerce operating experience each complete one run spanning 365 simulated days through the human operations dashboard over five calendar days. We report the mean across three LLM or Rule-based runs or across the three human participants.
Store Configuration.
Each agent operates a store initialized with a cash balance of RMB 2,000 and a security deposit of RMB 1,000 and supports at most 50 active listings. The fixed fines are RMB 8 for Return and Refund, RMB 5 for Bad Review, Stockout, or insufficient balance, and RMB 3 for Late Shipment. Cancellation and Returnless Refund incur no additional fine. These fine settings are consistent with the corresponding real-world platform rules.
Context Management.
To support interaction over the full 365 day horizon, both frameworks compress long interaction histories. When a ReAct history reaches 160,000 tokens, the evaluated model receives a reminder to summarize important information into persistent memory before the history is truncated to the most recent 30,000 tokens. Hermes uses its default context summarization procedure and sets the summary model to the same evaluated model.
Evaluation Metrics.
Net Profit Margin is terminal net profit divided by GMV. Average Store Rating and Average Active Listings are daily means of the published store rating and active listing count. Order Anomaly Rate is the share of generated orders affected by at least one cancellation, stockout, insufficient balance failure, late shipment, refund, or bad review. Sustained Window Rate is the minimum share of scheduled decision windows containing at least one environment tool call across all rolling 30 day periods. Total Tool Calls excludes calls that end a decision window and all framework internal tools.
Demand and Supplier Configuration.
Listing exposure uses , days, , and . Supplier inventory is initialized at 20 to 399 units, with capacities of 50 to 499 units and hourly replenishment of 1 to 19 units. Baseline supplier dispatch and logistics times span 1 to 47 and 12 to 72 hours. Supplier abnormalities recover after 168 to 672 hours, Price Change applies a factor sampled from 0.9 to 1.5, and Shipment Delay adds 12 to 96 hours.
Order Lifecycle Configuration.
The default promised shipment time is 48 hours. The transition from Delivered to Settled and the realization of outcomes after delivery are each sampled within 168 hours.
Rating Configuration.
The outcome scores are , , , , , and , and the evidence weights are , , , , , and , ordered as a normal transition to Settled, Late Shipment, Return and Refund, Returnless Refund, Bad Review, and Stockout. Store ratings use , , and . The rating thresholds , , , and map to demand multipliers , , , , and across the five resulting intervals. Diagnostic listing ratings use the same initial rating and prior weight with a 90 day evidence half life.
Appendix K Tool Sets and Skills
ReAct uses only the 26 MerchantBench tools through which the agent accesses observable fields and controls the store. Table 2 lists the complete MerchantBench tool set. Among them, the get_daily_report tool returns the report published for the current simulation date and states that its evidence is current only through the previous date. Table 3 gives an English translation of the Daily Market Report for June 10, 2025. Hermes provides the same MerchantBench tools together with the built in tools and skills listed in Tables 5 and 6 (17).
Appendix L Agent Inputs and Observable Fields
The environment gives each evaluated agent one system prompt at registration and a compact observation at every 12 hour decision window. Both frameworks receive the shared MerchantBench task prompt in Table 9 and access the same 26 MerchantBench tools. For ReAct, the MerchantBench task prompt is the complete system prompt and the 26 tools are the entire action space. Hermes wraps the same task prompt within its official runtime template, shown in Table 11, and augments the 26 MerchantBench tools with its built in tools and skills. Table 18 gives a representative observation containing order transitions, an upstream price change, current finances, and the shop rating.
MerchantBench is partially observable because the merchant interface returns current public and realized operational evidence while retaining future demand, event hazards, and presampled outcomes inside the simulator. Tables 19 to 21 summarize this boundary. Visible fields are returned directly by at least one merchant tool. Some visible order fields remain empty until the corresponding lifecycle event is realized. Hidden fields are never returned through merchant tools.
| Tool | Access | Description |
| Product Sourcing | ||
| get_daily_report | Read | Returns the daily market report for the current simulation date with market news and opportunity signals |
| search_products | Read | Searches the visible Product Catalog using public fields |
| get_product_detail | Read | Returns visible product, logistics, rating, and supplier fields |
| get_supplier_profile | Read | Returns the public supplier profile and visible product count |
| list_supplier_products | Read | Lists the currently visible products from one supplier |
| Listing and Pricing Control | ||
| list_product | Write | Adds products to the store at specified selling prices |
| delist_product | Write | Removes products from the store |
| adjust_price | Write | Changes selling prices for active listings |
| review_my_listings | Read | Reviews listing age, sales velocity, fines, and fulfillment backlog |
| query_my_listings | Read | Returns current listings with cumulative sales, profit, and fines |
| query_store_performance | Read | Summarizes store outcomes by day or week |
| query_product_sales_stats | Read | Ranks product outcomes and reports abnormality counts |
| Cash-Flow Management | ||
| query_balance | Read | Returns the cash balance, security deposit, funds in transit, receivables, and fines |
| get_store_snapshot | Read | Summarizes orders, supply, cash, listings, and store rating |
| query_platform_rules | Read | Returns capital, settlement, penalty, and closure rules |
| query_cash_pipeline | Read | Summarizes receivable aging and active order cost exposure |
| Supplier and Order Monitoring | ||
| query_supply_chain_anomalies | Read | Returns new or current supplier abnormalities and affected listings |
| query_my_orders | Read | Searches historical orders with logistics and accounting fields |
| query_open_orders | Read | Returns active orders with fulfillment timing and economics |
| query_order_updates | Read | Returns status changes since the previous observation window |
| query_order_detail | Read | Returns one order’s full status timeline, accounting, and penalties |
| Agent Support and Control | ||
| read_memory_doc | Read | Reads the run local agent memory document |
| write_memory_doc | Write | Replaces the run local agent memory document |
| get_observation | Read | Returns the current rendered observation |
| list_tools | Read | Returns tool schemas after scenario filtering |
| end_of_step | Control | Releases the current decision window |
| Tool | Access | Description |
| Execution and Files | ||
| terminal | Execute | Executes shell commands in a persistent environment |
| process | Manage | Monitors and controls background processes |
| execute_code | Execute | Runs Python programs that call Hermes tools and process their outputs |
| read_file | Read | Reads text files with line numbers and pagination |
| write_file | Write | Creates or replaces files and checks supported formats |
| patch | Write | Applies targeted file edits and returns a unified diff |
| search_files | Read | Searches file names and contents |
| Memory and Skills | ||
| memory | Write | Stores durable facts that persist across sessions |
| session_search | Read | Searches messages from previous Hermes sessions |
| skills_list | Read | Lists available skills and their descriptions |
| skill_view | Read | Loads skill instructions and linked resources |
| skill_manage | Write | Creates, revises, or deletes skills |
| Planning and Coordination | ||
| todo | Manage | Maintains the task list for the current session |
| clarify | Interact | Requests clarification, feedback, or a decision from the user |
| delegate_task | Delegate | Assigns independent tasks to subagents |
| Projects and Output | ||
| project_list | Read | Lists available project workspaces |
| project_create | Write | Creates and activates a project workspace |
| project_switch | Write | Switches the active project workspace |
| text_to_speech | Generate | Converts text into speech audio |
| image_generate | Generate | Generates or edits images from prompts and references |
| Skill | Category | Description |
|---|---|---|
| apple-notes | Apple | Creates, searches, and edits Apple Notes |
| apple-reminders | Apple | Adds, lists, and completes Apple Reminders |
| findmy | Apple | Tracks Apple devices and AirTags |
| imessage | Apple | Sends and receives iMessages and SMS |
| claude-code | Autonomous agents | Delegates coding tasks to Claude Code |
| codex | Autonomous agents | Delegates coding tasks to OpenAI Codex |
| hermes-agent | Autonomous agents | Configures and extends the Hermes Agent codebase |
| opencode | Autonomous agents | Delegates coding and review tasks to OpenCode |
| computer-use | General | Operates desktop interfaces through visual interaction |
| architecture-diagram | Creative | Creates architecture and infrastructure diagrams |
| ascii-art | Creative | Generates and transforms ASCII art |
| ascii-video | Creative | Converts video and audio into ASCII video |
| baoyu-infographic | Creative | Produces infographics using reusable layouts and styles |
| claude-design | Creative | Designs standalone HTML artifacts |
| comfyui | Creative | Generates images, video, and audio with ComfyUI |
| design-md | Creative | Authors and validates DESIGN.md specifications |
| excalidraw | Creative | Creates hand drawn Excalidraw diagrams |
| humanizer | Creative | Revises text to remove formulaic AI phrasing |
| manim-video | Creative | Produces mathematical and algorithmic animations |
| p5js | Creative | Creates interactive p5.js sketches and generative art |
| popular-web-designs | Creative | Applies established web interface design systems |
| pretext | Creative | Supports interactive creative browser demonstrations |
| sketch | Creative | Produces alternative HTML interface mockups |
| songwriting-and-ai-music | Creative | Supports songwriting and AI music prompting |
| touchdesigner-mcp | Creative | Controls TouchDesigner through an MCP interface |
| Skill | Category | Description |
|---|---|---|
| jupyter-live-kernel | Data science | Performs iterative analysis in a persistent Jupyter kernel |
| dogfood | General | Conducts exploratory testing of web applications |
| himalaya | Manages email through the Himalaya command line interface | |
| codebase-inspection | GitHub | Measures codebase size, languages, and composition |
| github-auth | GitHub | Configures tokens, keys, and command line authentication |
| github-code-review | GitHub | Reviews pull request diffs and inline comments |
| github-issues | GitHub | Creates and manages GitHub issues |
| github-pr-workflow | GitHub | Manages branches, commits, checks, and pull requests |
| github-repo-management | GitHub | Clones, creates, forks, and maintains repositories |
| gif-search | Media | Searches and downloads GIF content |
| heartmula | Media | Generates songs from lyrics and style tags |
| songsee | Media | Extracts and visualizes audio features |
| youtube-content | Media | Converts YouTube transcripts into written content |
| huggingface-hub | MLOps | Searches, downloads, and uploads models and datasets |
| evaluating-llms-harness | MLOps | Evaluates language models with standard benchmarks |
| weights-and-biases | MLOps | Tracks experiments, sweeps, and model artifacts |
| llama-cpp | MLOps | Runs local GGUF model inference |
| serving-llms-vllm | MLOps | Serves language models with vLLM |
| audiocraft-audio-generation | MLOps | Generates music and sound with AudioCraft |
| segment-anything-model | MLOps | Performs prompt based image segmentation |
| obsidian | Note taking | Reads, searches, creates, and edits Obsidian notes |
| Skill | Category | Description |
|---|---|---|
| airtable | Productivity | Manages Airtable records and queries |
| google-workspace | Productivity | Operates Gmail, Calendar, Drive, Docs, and Sheets |
| maps | Productivity | Provides geocoding, points of interest, routes, and time zones |
| nano-pdf | Productivity | Edits PDF text and document metadata |
| notion | Productivity | Manages Notion pages and databases |
| ocr-and-documents | Productivity | Extracts text from PDFs and scanned documents |
| petdex | Productivity | Installs and selects animated Hermes mascots |
| powerpoint | Productivity | Creates and edits presentation decks |
| teams-meeting-pipeline | Productivity | Operates the Teams meeting summary pipeline |
| arxiv | Research | Searches arXiv by topic, author, category, or identifier |
| blogwatcher | Research | Monitors blogs and syndicated feeds |
| llm-wiki | Research | Builds and queries an interlinked knowledge base |
| polymarket | Research | Queries prediction markets, prices, and order books |
| research-paper-writing | Research | Supports machine learning paper development and submission |
| openhue | Smart home | Controls Philips Hue lights, rooms, and scenes |
| xurl | Social media | Reads and operates X through its command line interface |
| hermes-agent-skill-authoring | Software development | Authors and validates Hermes skill packages |
| node-inspect-debugger | Software development | Debugs Node.js through the inspector protocol |
| plan | Software development | Produces actionable implementation plans |
| python-debugpy | Software development | Debugs Python with pdb and debugpy |
| requesting-code-review | Software development | Performs structured review before integration |
| simplify-code | Software development | Refines recent code changes with parallel review |
| spike | Software development | Runs disposable experiments before implementation |
| systematic-debugging | Software development | Applies a structured root cause debugging process |
| test-driven-development | Software development | Applies test driven development workflows |
| yuanbao | General | Operates Yuanbao groups and member queries |
| Product field | Access | Meaning |
| product_id | Visible | Stable identifier for a Product in the Product Catalog |
| name | Visible | Marketplace product title used for retrieval and comparison |
| category | Visible | One of the ten normalized first level product categories |
| quantity | Visible | Current effective supplier inventory after replenishment |
| price | Visible | Current procurement price offered by the supplier |
| historical_avg_rating | Visible | Historical product rating obtained from the source platform |
| logistics_hours | Visible | Baseline transit time from supplier dispatch to delivery |
| is_listed_by_supplier | Visible | Current procurement availability, exposed as supplier_available |
| ref_price | Hidden | Reference price used in the price response term of the demand model |
| base_price | Hidden | Supplier price restored after a temporary Price Change ends |
| cancel_rate | Hidden | Product level probability used to sample Cancellation |
| refund_rate | Hidden | Product level probability used to sample Return and Refund |
| only_refund_rate | Hidden | Product level probability used to sample Returnless Refund |
| bad_review_rate | Hidden | Product level probability used to sample Bad Review |
| max_quantity | Hidden | Inventory capacity used by the supplier replenishment process |
| hourly_increment | Hidden | Hourly supplier inventory replenishment amount |
| elasticity | Hidden | Product specific price elasticity used by the demand model |
| market_curve | Hidden | Real-world product level demand history over 365 days |
| quantity_updated_t | Hidden | Internal timestamp used for lazy inventory replenishment |
| price_recover_t | Hidden | Prescheduled end time of an active Price Change |
| delist_recover_t | Hidden | Prescheduled end time of an active Product Delisting |
| Supplier field | Access | Meaning |
| supplier_id | Visible | Stable supplier identifier |
| supplier_name | Visible | Public supplier name |
| shop_rating | Visible | Public supplier rating shared by all Products from the supplier |
| return_buyer_rate | Visible | Public repeat buyer rate returned by the supplier profile |
| supplier_age_years | Visible | Public supplier tenure in years |
| product_count | Visible | Number of currently available Products from the supplier |
| supplier_ship_hours | Visible | Current dispatch time for a Product from this supplier |
| base_ship_hours | Hidden | Dispatch time restored after a Shipment Delay ends |
| timeout_rate | Hidden | Product level hazard for Shipment Delay |
| price_change_rate | Hidden | Product level hazard for Price Change |
| supplier_delist_rate | Hidden | Product level hazard for Product Delisting |
| timeout_active | Hidden | Internal indicator of an active Shipment Delay |
| timeout_recover_t | Hidden | Prescheduled end time of an active Shipment Delay |
| Order field | Access | Meaning |
| order_id | Visible | Stable identifier for an individual customer order |
| product_id, product_name | Visible | Product identity associated with the order |
| supplier_id, supplier_name | Visible | Supplier identity associated with the order |
| order_time | Visible | Calendar and simulation time at which the order was placed |
| current_status | Visible | Latest realized lifecycle state |
| status_age_hours | Visible | Elapsed time since the latest realized status transition |
| expected_delivery_time | Visible | Current delivery estimate computed from realized timing information |
| delivered_time | Visible | Merchant facing delivery timestamp populated after delivery |
| late_time | Visible | Merchant facing timestamp populated only after Late Shipment is realized |
| sale_price | Visible | Merchant selling price recorded when the order was created |
| purchase_price | Visible | Procurement price recorded when the order was created |
| supplier_ship_hours | Visible | Supplier dispatch duration recorded for the order |
| supplier_logistics_hours | Visible | Baseline post dispatch logistics duration |
| actual_logistics_hours | Visible | Realized transit duration populated after delivery |
| realized_revenue | Visible | Revenue credited from outcomes realized so far |
| realized_cost | Visible | Procurement cost realized so far |
| total_penalty | Visible | Sum of penalties already applied to the order |
| net_profit | Visible | Realized revenue minus realized cost and total penalty |
| profit_finalized | Visible | Indicator that no further profit component remains unresolved |
| status_log | Visible | Realized sequence of lifecycle states and their timestamps |
| preset_anomaly | Hidden | Presampled future outcome among normal fulfillment and four customer abnormalities |
| preset_anomaly_t | Hidden | Internal realization time of the presampled abnormal outcome |
| settlement_delay_steps | Hidden | Presampled delay from delivery to final settlement |
| purchase_t, shipped_t, delivered_t, settled_t | Hidden | Raw internal transition times, with only realized merchant facing views exposed |