NandhaKishorM/laya leads the week with a non-autoregressive decision engine. Hand it text, an email or a JSON document plus a typed question, and one 33-millisecond forward pass returns a choice, a score or a yes/no flag, with nothing generated and nothing to parse. Three checkpoints cover English, more than one hundred languages and typed workflows, and a built-in router picks one per request. jaredpalmer/kev pushes the same idea toward open weights: 0.8B, 4B and 9B models on Qwen3.5 that answer choice, yes/no and rating questions with probabilities instead of a single label, which you can run or train yourself on CUDA, ROCm or Apple Silicon. Apple Silicon users can also try mizorewww/laya-mlx, a native MLX port of laya that answers a short typed question in about 13 milliseconds with zero output tokens.

Memory and orchestration are covered by two very different projects. vectorize-io/hindsight treats agent memory as three operations, retain, recall and reflect, organizes what an agent knows into memory banks and mental models, and connects to coding agents through an MCP server. google/ax, a declarative orchestrator from Google in the spirit of Kubernetes, has one YAML file declare the task, pre-wired Git repos and MCP servers, a network allowlist and the model; ax apply launches it into a sandbox with CPU and memory limits, and ax ssh opens a shell inside a running agent. If hindsight answers what an agent knows, ax answers where it runs.

Several entries let agents operate real software. paperclipai/paperclip runs a team of agents like a company: tasks arrive as tickets with approval gates you verify from diffs, screenshots and tests, an org chart assigns roles and scoped secrets, heartbeats wake agents on a schedule, and a monthly budget per agent caps token spend. browser-use/jev-ultrafast drives a browser by rendering each page as a numbered table of controls, so a single model call picks both the operation and the target element. jev-chat/jev-chat-jarvis brings a copilot to an Android phone: it reads on-screen chats through the accessibility service, judges the other side’s intent and drafts ranked replies, but it never hooks the app and never presses send for you.

For infrastructure and local tooling, hydra-db/hydradb keeps a Rust graph database on S3-compatible object storage and queries pinned snapshots in OpenCypher, while rocketride-org/rocketride-server builds portable AI pipelines on a VS Code canvas and runs them with a multithreaded C++ engine on your own hardware. debpalash/VoiceStudio keeps voice design, transcription and dubbing inside one local Electron app.

Not everything this week is for agents. eternity4719/HowToLiveBetter collects hundreds of everyday tips that each state the cost, the return and the strength of the evidence, citing only journal papers and official documents. rohitg00/ai-engineering-from-scratch is a free curriculum where every lesson ends in a reusable artifact, and mattpocock/skills makes your coding assistant ask the right questions before it writes code. DietrichGebert/ponytail exists to stop an agent turning a small fix into a new framework, and if you alternate between several coding tools, farion1231/cc-switch manages their model providers and MCP servers from one desktop panel.