AI X-feeddaily signal from hand-vetted sources

2026-09-18

52 signal posts

Relevance 7/10research

ScientistTwo screens research ideas on subsets, runs full experiments and ablations, then revises work using review feedback.

Its staged screening and attribution loop offers patterns for making autonomous agent experiments more reliable.

@dair_ai · 2026-09-18 · autonomous-agents, research, experimentation, evaluation

Relevance 8/10project_demo

Roboclaw runs a team server, joins Discord, and answers questions about context across active and past sessions.

Cross-session context lookup is a practical pattern for giving teams continuity across agent work and meetings.

@steipete · 2026-09-18 · openclaw, agent-ops, session-context, discord

Relevance 8/10technique

A teammate cleans up sessions before PRs land; you can ask OpenClaw from its home sidebar to reorganize sessions.

Session cleanup can keep long-running agent work easier to navigate and review.

@steipete · 2026-09-18 · openclaw, agent-workflow, sessions, code-review

Relevance 8/10tool_release

Computer-use actions work across Crabbox OSes, giving agents an option beyond screenshot-based interaction.

More direct computer interaction can make cross-platform agent workflows more efficient than screenshot loops.

@steipete · 2026-09-18 · openclaw, computer-use, agents

Relevance 8/10tool_release

Crabbox boxes support Linux, macOS, and Windows.

Cross-platform boxes make the same agent testing workflow usable across different development environments.

@steipete · 2026-09-18 · openclaw, sandboxing, cross-platform

Relevance 9/10tool_release

OpenClaw can move a local web app into Crabbox and show it through VNC or a portal when testing needs a box.

Lets an agent shift from local work to an isolated test environment without restarting the workflow.

@steipete · 2026-09-18 · openclaw, sandboxing, browser-agents, remote-desktop

Relevance 5/10project_demo

Links to a featured GPT-6 Astra build as an example of what developers could ship.

It offers a concrete project to inspect for inspiration, but no implementation lessons in the post.

@OpenAIDevs · 2026-09-18 · 3d, gpt, generative-ai

Relevance 5/10project_demo

Showcases community 3D builds made with GPT-6 Astra.

The gallery may spark ideas for multimodal projects, though the post offers no build details.

@OpenAIDevs · 2026-09-18 · 3d, gpt, generative-ai

Relevance 7/10technique

Suggests semantic triggers and semantic crons as alternatives to purely time-based automation.

This framing could help make OpenClaw tasks respond to relevant events instead of fixed schedules.

@hwchase17 · 2026-09-18 · agents, automation, scheduling

Relevance 5/10research

Points to a thread about information recoverable from unexpected signals, including CPU temperature fluctuations.

It may broaden how you think about unintended information channels in computing systems.

@emollick · 2026-09-18 · side-channels, information-theory, cpu

Relevance 5/10news

LangChain’s Sydney Runkle is set to present a deep dive on using Jev for agents next week.

The session could offer practical examples of applying Jev in agent workflows.

@hwchase17 · 2026-09-18 · agents, jev, event

Relevance 9/10research

A 176-setting study finds rule-based elision before LLM summarization is cost-effective; context management mainly prevents overflow failure

Its tested tradeoffs can help you simplify harnesses and choose context, planning, and tool strategies for different models.

@dair_ai · 2026-09-18 · coding-agents, context-engineering, harness-design, evaluation

Relevance 9/10research

NVIDIA’s SoL-Pi auto-evolves coding-agent harnesses; four mechanisms cut token traffic nearly in half while matching baseline performance.

The mechanisms and reported cost savings give you concrete ideas to test in Claude Code or OpenClaw harnesses.

@omarsar0 · 2026-09-18 · agent-harness, context-engineering, coding-agents, cost-optimization

Relevance 8/10project_demo

A demo shows HumanLayer handling large, design-heavy engineering tasks while OpenClaw and AmpCode orbs handle smaller work.

The division of labor offers a concrete workflow to compare with your OpenClaw and coding-agent setup.

@dexhorthy · 2026-09-18 · agent-workflows, openclaw, humanlayer, coding-agents

Relevance 5/10news

Anthropic and Accenture plan to invest at least $1B each over five years in independent frontier-AI evaluation.

Signals growing investment in evaluation, a field relevant to building and assessing reliable agents.

@AnthropicAI · 2026-09-18 · ai-evaluation, anthropic, industry

Relevance 5/10opinion

A game’s author says deciding what to leave out took more effort than writing the shipped code.

A useful reminder that scope reduction can take more judgment than implementation in your own projects.

@badlogicgames · 2026-09-18 · product-design, scope

Relevance 5/10technique

Shares a guide to React aimed at engineers moving into frontend work.

A practical learning resource if your team is adopting React and you are new to frontend development.

@badlogicgames · 2026-09-18 · react, frontend, learning

Relevance 8/10tool_release

Points to Claude Code's new mods system and a repository of example mods.

The examples give you starting points for extending the Claude Code workflows you use daily.

@simonw · 2026-09-18 · claude-code, extensions, agents

Relevance 8/10opinion

Argues that solving a problem yourself builds understanding an agent may miss, so delegated solutions can be wrong or shallow.

A useful reminder to keep human problem-discovery in the loop when delegating implementation to agents.

@badlogicgames · 2026-09-18 · agents, vibe-coding, software-engineering

Relevance 8/10news

Claude Code's AGENTS.md support removes the need for CLAUDE.md files that only import AGENTS.md.

Use one shared AGENTS.md instruction file instead of maintaining a Claude-only wrapper.

@simonw · 2026-09-18 · claude-code, agents-md, instructions

Relevance 8/10research

A paper investigates how harness design affects coding-agent performance.

Its findings could inform harness choices for Claude Code and your own agents.

@_akhaliq · 2026-09-18 · coding-agents, harness, agent-evaluation

Relevance 6/10opinion

Argues developers should express intent at a higher level and let a compiler handle low-level optimization.

A useful design principle for agent interfaces: specify intent instead of hand-authoring execution details.

@lateinteraction · 2026-09-18 · abstraction, compilers, software-design

Relevance 8/10technique

A chess test suggests routing easy classifications to fast models and reserving reasoning models for harder calls.

Route easy cases to cheap models and reserve reasoning for ambiguous decisions to lower pipeline cost.

@nutlope · 2026-09-18 · multi-model, classification, routing, cost-optimization

Relevance 6/10opinion

Emollick argues today's AI capability overhang leaves years of underused potential and outlines four ways to close it.

A useful frame for finding higher-leverage AI workflows before assuming model limits.

@emollick · 2026-09-18 · ai-adoption, workflows, productivity

Relevance 8/10tool_release

Anthropic has published the built-in AGENTS.md mod source in the Claude Code repository.

The implementation offers a concrete starting point for custom Claude Code harness mods.

@trq212 · 2026-09-18 · claude-code, agents-md, harness, github

Relevance 9/10tool_release

Claude Code mods will power AGENTS.md support and let projects customize instruction handling.

You can extend the Claude Code harness with project-specific instruction behavior beyond a static AGENTS.md.

@trq212 · 2026-09-18 · claude-code, agents-md, harness, customization

Relevance 10/10tool_release

Claude Code 2.1.277 can use AGENTS.md when a folder has no CLAUDE.md; toggle it in /config.

Lets you reuse AGENTS.md instructions in Claude Code without renaming files.

@trq212 · 2026-09-18 · claude-code, agents-md, agent-tooling

Relevance 8/10technique

Use agents to help, but inspect data yourself; the linked post covers auto-evals.

Combining agent assistance with direct data review helps catch issues automated evals can miss.

@HamelHusain · 2026-09-18 · agents, evals, data

Relevance 9/10technique

For large trace reviews, find the first upstream failure and make supporting evidence expandable.

A practical way to make agent and eval trace reviews faster without hiding important evidence.

@HamelHusain · 2026-09-18 · evals, debugging, traces

Relevance 8/10tool_release

The gog CLI now offers an MCP server for connecting Google-service workflows to agent clients.

It gives an OpenClaw or other MCP setup a direct path to Google-service tools.

@steipete · 2026-09-18 · mcp, google, cli, integrations

Relevance 6/10project_demo

A shared agent session lets collaborators see each other's typing and avoid sending duplicate instructions.

The shared view can reduce conflicting or duplicated prompts when people work with an agent together.

@steipete · 2026-09-18 · agents, collaboration, workflow

Relevance 9/10research

EvoSkill v2 turns failed runs into persistent skills; separating skill-writing from grading helped spreadsheet success rise from 3/120 to 21

Separating skill-writing from grading and adding human review helps agents learn from failures without reward hacking.

@omarsar0 · 2026-09-18 · agent-skills, self-improvement, evaluation, agents

Relevance 6/10tool_release

Multi-account support is now available across most plugins, covering personal, work, and project accounts.

Separate accounts can make plugin-based workflows easier to use across personal and work contexts.

@OpenAIDevs · 2026-09-18 · plugins, multi-account, developer-tools

Relevance 5/10tool_release

Jev returns calibrated probabilities for defined questions rather than generating natural-language answers.

A distinct model interface could suit classification tasks, though no integration details are given.

@altryne · 2026-09-18 · probability-models, ai-tools, llms

Relevance 5/10tool_release

Jev charges $42 per billion input tokens, with no charge for its probability outputs.

The pricing model is useful context, but the post offers little for agent-building workflows.

@altryne · 2026-09-18 · pricing, probability-models, ai-tools

Relevance 7/10tool_release

AgentCloak replaces personal details in prompts with stand-ins, then restores them in AI responses.

Could reduce exposure of sensitive data when using browser-based AI for logs or support tickets.

@omarsar0 · 2026-09-18 · privacy, browser-extension, llm-tools

Relevance 5/10opinion

Speculates that Jev could build a load balancer that beats some heuristics when speed demands are modest.

Suggests a concrete agent-built systems experiment, but gives no results or implementation details.

@thorstenball · 2026-09-18 · agents, load-balancing, experimentation

Relevance 8/10project_demo

A multi-agent presentation workflow used encouraging quality prompts while leaving the final product to Claude.

Try guiding agents with quality questions during work instead of taking over the finished output.

@emollick · 2026-09-18 · agents, prompting, presentations, workflow

Relevance 5/10opinion

Suggests creating directories and projects on a runner directly from its picker.

Highlights a handy runner workflow idea for reducing friction when starting agent tasks.

@thorstenball · 2026-09-18 · developer-tools, agents, workflow

Relevance 5/10news

Anthropic and Accenture plan to invest at least $1B each in frontier AI evaluation over five years.

Signals growing investment in independent evaluation, though the post offers no methods to apply.

anthropic.com · 2026-09-18 · anthropic, evaluation, ai-safety

Relevance 7/10tool_release

Herdr is described as a tmux-like way to manage agents.

A tool to explore for running and coordinating agents in your development setup.

@GeoffreyHuntley · 2026-09-18 · agents, developer-tools, workflow

Relevance 8/10technique

Shares a deep dive on why software factories fail, reading agent-written code, and designing a software factory.

The practical guidance can inform how you build and review agent-driven coding workflows.

@dexhorthy · 2026-09-18 · coding-agents, software-factory, code-review

Relevance 6/10opinion

Argues that optimizing for scale before product-market fit is premature, especially as companies reassess fit frequently.

The reminder to validate demand before overengineering applies to AI-built products too.

@thorstenball · 2026-09-18 · ai-coding, product-market-fit, software-engineering

Relevance 7/10tool_release

Jev adds repeatable directory discovery paths and a flag to set search depth.

These options make it easier to tune an agent tool's project-file discovery.

@thorstenball · 2026-09-18 · agents, cli, tooling

Relevance 7/10project_demo

Shares the Jev shell-history integration code, which the author says they never inspected.

The repository offers code to examine or adapt for shell-aware agent workflows.

@thorstenball · 2026-09-18 · agents, shell, github

Relevance 6/10project_demo

Shows Jev choosing the next command from shell history.

A glimpse of agent-assisted command selection may spark useful terminal workflow ideas.

@thorstenball · 2026-09-18 · agents, shell, developer-tools

Relevance 8/10opinion

Proposes replacing some MCP integrations with codemode, OpenAPI, and RAG over API docs.

This gives builders a concrete alternative to compare when designing tool access for agents.

@mitsuhiko · 2026-09-18 · mcp, openapi, agents, rag

Relevance 6/10news

Announces a talk on 12-factor factories, slopcodebench, and the Gas Town situation at AGNTCon + MCPCon.

The talk may provide transferable agent-factory ideas and useful MCP ecosystem context.

@dexhorthy · 2026-09-18 · agents, mcp, agent-architecture

Relevance 6/10news

Promotes a podcast episode on Jev primitives, assistants, computer use, and AI discussion pacing.

The episode could offer practical ideas on emerging agent and computer-use workflows.

@altryne · 2026-09-18 · jev, agents, computer-use

Relevance 5/10opinion

Argues that rewriting dependencies can limit security blast radius, while reducing outside scrutiny of the code.

It highlights a real tradeoff to weigh when choosing dependencies for personal agent infrastructure.

@GeoffreyHuntley · 2026-09-18 · security, dependencies, software-engineering

Relevance 8/10technique

Codex can request a blank context and have the LLM maintain notes, rather than relying on conventional compaction.

This expands the design space for managing long-running agent state and context limits.

@mitsuhiko · 2026-09-18 · agents, compaction, context-engineering

Relevance 7/10opinion

Suggests Jev could make the pruning harnesses already do during compaction more cost-effective.

It offers a practical angle for reducing agent context costs without assuming compaction can be skipped.

@mitsuhiko · 2026-09-18 · agents, compaction, context-engineering

Curated one-line summaries; every title opens the original post. Selected and summarized automatically from hand-vetted sources by a pipeline running on a Raspberry Pi. Numbers are relevance scores (0–10) assigned by the curator model against an applied-AI rubric. Times are US Eastern. Updated every 4 hours.