commerce-agents is Anthropic's official Apache-2.0 reference blueprint (created 2026-09-01, 289+ stars in 2 days) for building two Claude agents — a customer-facing shopping agent and a staff-facing merchant agent. Each agent is defined once (prompt, skills, tool contracts, gates) and runs on the Messages API, the Claude Agent SDK, or Managed Agents, with four runnable verticals (retail, travel, telecom, entertainment) plus a Claude Code plugin that scaffolds agents against your own systems. This review covers the two-agent architecture, the five flows of each, the safety model (fencing, provenance gates, staged merchant writes, approval surfaces), the honest limitations (reference-only, not maintained, no contributions accepted), and who it's for.
Agent Frameworks
3 tools reviewed
DeepSeek Harness (dsh) is DeepSeek's open-source agent harness where every capability — models, tools, skills, sessions, sandboxes, storage, loops, scheduling, UI — is a swappable plugin. Built on the Cordis meta-framework. We ran it locally: npx @deepseek-ai/dsh web serves a web UI at 127.0.0.1:3080. Review covers the plugin architecture, the every-run-is-traceable session log, the 39.5k-star launch, and what HN developers think of the TypeScript/Cordis design choices.
reef is Apache-2.0 continual-learning infrastructure from Human-Agent-Society (created 2026-08-31, 270+ stars in four days, 10+ active contributors) that turns agent inference into a closed learning loop. You install an agent harness the way you install codex — curl a Reef endpoint — and point its model requests at Reef's OpenAI- and Anthropic-compatible inference API instead of the provider's. Behind the scenes Reef runs a four-step cycle: Serve (record every interaction), Observe (match scored feedback to receipts), Grow (build candidate updates from eligible records), Commit (evaluate and publish winners into a versioned artifact history). Two learning surfaces are supported: model weights (SGLang + slime training) and agent harnesses (rules, skills, prompts, config — evolved via cordis, no GPU required for the skill-pool recipe). Cookbook recipes include SAO for verifier-scored task streams, OpenClawRL for agent traffic without explicit reports, TTTD for repeated scored attempts, and SkillClaw for evolving an agent's skill pool. This review covers the four-step loop, the receipt/report API, the harness-install flow, the honest caveats (4 days old, GPU-class training recipes, feedback plumbing is on you), and who it's for.