Scout Briefing — Tuesday, August 25, 2026¶
🧭 Today's Thesis¶
Agent skills are becoming software dependencies before the ecosystem has learned to evaluate them like software dependencies. Catalogs, mirrors, plugin formats, and app servers have solved discovery and distribution; ACES-style paired runtime trials expose how little install counts and structural scans say about actual value. The next durable control plane will admit a capability only when identity, permissions, task lift, version, outcome, and rollback travel together.
Coverage & methodology
Evidence and velocity provenance: The exact pre-collected run directory was reused and no collector was rerun. GitHub, GitHub Search, and arXiv were healthy; HN was stale and repaired through live discovery. Fifteen of 17 direct-source URLs hydrated successfully; the OneCLI and coding-identity HN pages returned HTTP 429, so verification of those two discussions is explicitly limited to bounded pre-collected records. The optional YouTube lane supplied two transcript-verified supporting items and did not control any release, benchmark, security, or adoption claim. GitHub daily, weekly, and monthly observations remain separate, and only
stars_todayaffected velocity or status.
🔥 Top Movers¶
- freestylefly/awesome-gpt-image-2 (2,449 ⭐ today, 15,693 total) — the raw board leader turns image prompts into versionable templates and skills, but it enters the ignore lane until fixed-seed quality and provenance evidence exists.
- openai/codex (1,994 ⭐ today, 117,191 total) — the largest on-lens runtime signal; its one-day total-star delta is directionally consistent, and status remains stable below the verified 4,159/day peak.
- stablyai/orca (982 reported today, 52,902 total) — cross-vendor fleet supervision stays hot, but a two-day registry gap means the total can be updated without a new acceleration or all-time-high claim.
- NousResearch/hermes-agent (896 today, 235,864 total) — clean acceleration from 454/day, with the one-day total delta of 809 supporting the direction of the measurement.
- Alishahryar1/free-claude-code (891 today, 49,072 total) and diegosouzapw/OmniRoute (667 today, 54,505 total) keep cost and provider routing visible; neither solves cost per accepted outcome.
- VoltAgent/awesome-agent-skills (602 today, 31,935 total), multica-ai/andrej-karpathy-skills (588, 206,589), and anthropics/claude-plugins-community (489, 1,387) make capability packaging the day's strongest cluster.
🎯 What Matters to Us This Week¶
- Skills finally have a useful deployment question: does this artifact add lift? ACES runs paired live trials with and without a skill under the same model, sandbox, workspace, task, and scorer, then reports Skill Lift. That is the right gate for the day's catalog and plugin surge: structural linting and star counts cannot show whether a skill improves completion, routing, tool use, or safety.
- MCP is becoming ordinary cross-language infrastructure while its hard problems move upward. The official MCP roadmap prioritizes agentic messaging, HTTP hardening, workload identity, delegated authority, progressive discovery, and SDK experience. The registry evidence is already broad: the TypeScript SDK showed about 34.6 million weekly npm downloads, the .NET SDK showed roughly 25.7 million total downloads, and Python SDK v2 is now the stable default. Connectivity is commodity; identity, permission, and lifecycle evidence are not.
- A team agent harness is becoming a security product, not just a workbench. The fresh OneCLI launch discussion describes isolated sandboxes, credentials kept out of model context, and deterministic approval. The unresolved issue is exact binding: approval must name the recipient, repository, operation, payload, expiry, and retry semantics, not merely “allow Gmail” or “allow this endpoint.” Deterministic hydration was rate-limited, so product details are treated as limited launch evidence rather than independently verified behavior.
- Review ownership is part of adoption cost. A direct coding-identity discussion asks how experienced developers retain craft and understanding as agents perform more implementation, while Argus frames QA capacity as the next bottleneck. The useful metric is accepted outcomes per reviewer-minute with replayable corrections—not generated lines, PR count, or agent starts.
🚀 What Changed the Frontier¶
- Interoperability moved from a command to a packaged host surface. OpenAI's August 24 release note deprecates
codex mcp-serverin favor of the Codex app server and directs Claude Code users to a Codex plugin. That is a small release with a large architectural tell: distribution and embedding now happen through app servers and plugin packages rather than a one-off protocol command. - Portable plugins reached mainstream developer surfaces. GitHub's August 10 Copilot release makes Agent Plugins 1.0 available across VS Code, CLI, SDK, and desktop app workflows. The next frontier is not another packaging format; it is evidence that permissions, versions, tests, and rollback survive movement between hosts.
- The harness became part of the object being evaluated. The Terminal Agents survey argues that realized behavior is jointly shaped by model, interface, harness, runtime, and environment and calls for replayable traces and process-level evidence. Model-only benchmarks are increasingly unable to explain why one delivery system succeeds and another fails.
🆕 First Appearances¶
- freestylefly/awesome-gpt-image-2 — first clean daily baseline at 2,449/day. Prompt-as-code is a real packaging pattern; adoption is deferred pending quality and provenance evaluation.
- VoltAgent/awesome-agent-skills — first baseline at 602/day. The catalog is useful for discovery, not as an install allowlist.
- multica-ai/andrej-karpathy-skills — first baseline at 588/day. A single instruction file is cheap to A/B test; its popularity is not task-lift evidence.
- anthropics/claude-plugins-community — first baseline at 489/day. The official read-only mirror makes community packages inspectable while leaving submission trust and revocation outside the repository.
🌱 Rising Stars¶
(Only consecutive, window-labelled daily observations are used.)
- NousResearch/hermes-agent — 454 → 896/day with an 809-star one-day total increase; a valid current acceleration signal without comparison to the legacy-unverified April peak.
- tinyhumansai/openhuman — 39 → 515/day with a 503-star total increase to 37,303; the first valid acceleration after its clean baseline.
- apache/maka — 51 → 411/day, 2,946 total. The current signal is rising below its verified 460/day peak, so this is not an ATH claim.
- can1357/oh-my-pi — reached 413/day and 27,185 total on a consecutive clean window; the operator value remains IDE-connected agent ergonomics, not raw runtime plurality.
📉 Fading¶
No new fading call crossed the required greater-than-80% drop from a verified daily peak. apple/coreai-models fell from 108 to 24/day, a 77.8% decline, so status remains stable rather than being rounded into a fade. Repositories seen only in weekly or monthly windows did not change velocity or status.
⚔️ Battles (same category, competing)¶
- Anthropic community plugins vs GitHub Agent Plugins 1.0 vs VoltAgent's catalog — Anthropic supplies a vendor-hosted mirror, GitHub supplies a cross-surface package format, and VoltAgent supplies cross-host discovery. None has yet combined verified publisher identity, least-privilege permissions, paired task lift, version pinning, and revocation into one portable receipt.
- Orca vs Proliferate vs OneCLI — Orca emphasizes bring-your-own-subscription fleet supervision; Proliferate emphasizes self-hosted cross-vendor coding workspaces; OneCLI emphasizes team sandboxes, credential mediation, and approvals. Compare accepted merges per reviewer-minute, isolation, exact-action grants, and exportable traces.
- Codex vs Hermes Agent vs OpenCode — Codex has the largest on-lens runtime signal, Hermes combines an open model family with its harness, and OpenCode supplies vendor-neutral TypeScript tooling. Terminal-agent research says the comparison must hold model, environment, task, and review policy constant before attributing the result to one component.
🔬 From Research¶
- Evaluating Skills, Not Just Agents — paired runtime trials measure whether a capability package adds value under fixed conditions; this is immediately actionable for plugin and skill intake.
- Terminal Agents — unifies terminal-mediated agency around model, interface, harness, runtime, environment, verification, and recovery rather than final answer alone.
- PrimeAgentOrchestrator — primes new coding-agent sessions from PostgreSQL and semantic-memory backends, a concrete implementation of cross-session context whose provenance and stale-memory failure modes should be tested.
- Nexus — retrieves and compresses tool signatures instead of prefilling every MCP schema, aligning with progressive discovery while making routing quality and cache trust first-class evaluation targets.
🔄 What's Changing¶
The ecosystem can now distribute a capability through a skill, plugin, app server, marketplace, or protocol SDK faster than a team can prove that the capability is useful and safe. That flips the bottleneck from authoring and connectivity to intake: identify the publisher, constrain authority, run paired tasks, preserve the trace, and revoke the package cleanly.
🧪 One Experiment Worth Running¶
- Paired plugin receipt test — choose one small TypeScript maintenance skill from the Anthropic mirror or VoltAgent catalog and ten fixed repository tasks. Run each task with and without the skill under the same model, sandbox, context, and CI; record completion, wrong-tool calls, reviewer corrections, security-policy violations, tokens, wall time, and accepted result. Expected upside: a low-cost allowlist gate. The decisive learning is whether the skill adds outcome lift after its review burden is counted.
⚠️ One Risk to Track¶
- A capability package gains broad authority before it proves lift. Trigger: a marketplace or host auto-installs or auto-updates a community plugin whose publisher, permission envelope, network reach, or version is not pinned. Downside: the package becomes an executable supply-chain dependency with agent-level credentials and no reproducible benefit. The current MCP security guidance reinforces audience binding, token isolation, PKCE, and confused-deputy controls; plugin hosts need an equivalent artifact-level contract.
🙅 One Thing to Ignore¶
- Skill and prompt catalogs ranked by stars.
awesome-gpt-image-2,awesome-agent-skills, andandrej-karpathy-skillsare strong attention signals and useful discovery leads, but catalog size, attribution, and stars do not measure safety or task lift. Revisit any individual artifact only after a paired task test, source/license review, permission inventory, and rollback check.
✍️ Writing Angle To Explore¶
- “Your agent skill is a dependency; where is its lockfile and A/B receipt?” — the timely tension is that plugin distribution is reaching IDE/CLI/app scale while ACES shows how to evaluate the artifact that packaging standards merely transport. The article note is saved in today's content exploration lane.
💡 Surprise Pick¶
apache/maka — not because it is another workspace, but because its append-only messages, tool calls, results, permissions, and termination events form the evidence substrate every plugin and fleet surface now needs. Its renewed 411/day signal is secondary; the useful test is whether an independent reviewer can reconstruct a consequential run without reading an unstructured chat transcript.
📊 Supply vs. Demand¶
| What's being built (supply) | What people want (demand) | Match? |
|---|---|---|
| Skill catalogs, community mirrors, portable plugin formats | Reusable behavior with verified lift, permissions, updates, and revocation | Weak — distribution is ahead of evidence |
| Cross-vendor fleet workbenches and sandboxed personal agents | Faster delivery without losing review ownership or exact authority | Partial — control surfaces exist; portable receipts do not |
| MCP SDKs across TypeScript, Python, and .NET | Stateless, secure, discoverable production integrations | Strong on connectivity; open on identity and delegation |
| Faster coding runtimes, local models, and quota routers | Maintainable output and explainable cost per accepted task | Weak — throughput and spend attribution remain disconnected |
| Agentic QA tools and security scanners | Regression coverage that keeps pace with generated changes | Partial — tools exist; shared workload evidence remains scarce |
| Reasoning-heavy local-agent settings | Reliable tool use with acceptable latency and cost | Narrowing — a recent community synthesis favors lower reasoning presets, but the evidence is community-reported |
📊 Category Pulse¶
| Category | New Today | Trending Count | Signal |
|---|---|---|---|
| Skills ecosystem | 4 | 10+ | 🔥 Distribution abundant; paired evaluation becomes urgent |
| Code dev tools | 0 true first appearances | 12+ | 🔥 Runtime competition shifts toward host and harness quality |
| Agent infra | 0 | 7+ | 📈 Sandboxes, grants, and append-only receipts converge |
| MCP tooling | 0 | 10+ plus 3 package registries | 📈 Cross-language commodity with identity work ahead |
| LLM eval/testing | 0 registered; 1 launch + 1 paper | 5+ | 📈 Skills and trajectories become regression targets |
| Memory/RAG | 0 | 5+ | ⚠️ Retrieval supply remains high; provenance and retention failures remain open |
Evidence and Catch-Up Notes¶
Open supporting detailSources, caveats, and catch-up notes
- The HN lane was the only required unhealthy lane. Fresh HN fallback evidence is present; one direct HN page hydrated successfully and two returned HTTP 429 with verification limitations recorded in their evidence metadata.
- The evidence archive contains 17 unique URLs across ten hosts; 15 hydrated successfully. It includes Reddit, official releases, security guidance, production/team evidence, npm, PyPI, NuGet, GitHub repositories, and two directly hydrated arXiv papers.
- The arXiv collector was healthy with 15 papers inside the seven-day research window, so no empty-array repair or rewrite was needed. The canonical research archive remains complete.
- The due
2026-W34.mdis already a completed August 17–23 seven-day synthesis with visible arXiv links, so it was retained rather than replaced. July's monthly market artifact already exists, so no monthly catch-up is due. - Tuesday is the first preferred content slot for W35; one exploration note is generated today. The Friday slot remains future work, not a missed catch-up.
- The raw snapshot retains 524 observations: 182 daily, 154 weekly, and 188 monthly. No metric was flattened into
stars_period.