Watchlist

title
Watchlist
type
overview
summary
Toolbox entries too early or too uncertain to commit to β€” re-check periodically for maturity, abandonment, or scope changes
tags
meta, watchlist
created
2026-04-24
updated
2026-09-14

Tools listed in index that aren't ready for actual use yet. Either too young (alpha, single-platform, single-author), too narrow (one-feature demo that may not grow), or too dependent on outside circumstances (a vendor's API, a regulatory window, a single corporate sponsor). The point of this page is to come back periodically and re-decide: graduated, abandoned, or still waiting.

How to use

Tag a toolbox entry with watchlist in its frontmatter and add a row here. Each row says what we're waiting on and when to re-check. When the verdict is in, remove the row and either drop the watchlist tag (graduated) or note the abandonment in the toolbox page itself (and keep the page β€” abandoned tools are still useful as references).

What "check later" means

Concrete signals to look at on each re-visit:

  • Activity β€” commits in the last 90 days, recent issues triaged, PRs merged
  • Star trajectory β€” flat-lining at the launch peak vs. continued growth
  • Release cadence β€” tagged versions or only main
  • Author bandwidth β€” solo project, employer-supported, or community-maintained?
  • Platform breadth β€” single-OS alpha vs. broadening
  • Stability claims β€” has the README dropped the "alpha / not for production" warning?
  • Competitor consolidation β€” has the obvious incumbent absorbed the idea or shipped a comparable feature?

A tool that fails most of these after 6–12 months is almost certainly dead and should be retired from the watchlist.

Current entries

Tool Watching for Next check Reason
codealmanac macOS-only (launchd-bound), requires Codex or Claude Code, YC S26 startup with a hosted-product path not yet visible, lifecycle agents run unattended with broad non-interactive filesystem access 2026-10-19 Transcript-harvesting plus a scheduled gardening pass is the right shape for a codebase wiki and the two ideas are worth stealing regardless. Want to see a Linux port, whether the unattended-agent trust model survives contact with users, whether the almanac/ boundary gets a real sandbox, and what the business model turns out to be.
openhuman Self-declared early beta with 170 open issues, enormous scope (memory + orchestration + workflows + meetings + messaging + voice + media gen + agent payments) in one app, nine-day GitHub trending run is launch-hype signal, GPL-3.0 unlike the rest of the toolbox 2026-10-19 Rust core with Privacy Mode enforced in-core plus plain-markdown memory makes the local-first claim more credible than most, and tinyagents/tinyflows being separately open-source means the durable parts may outlive the app. Want to see the scope narrow or the quality hold, issue count come down, contributors beyond the founding team, and whether Privacy Mode survives audit.
freeink freeink-sdk is seven weeks old (2026-06-03) with 38β˜… against CrossPoint's 6.2k; de-link has no license file and no push since 2026-04-21; the hardware-abstraction claim is unproven off the boards it was extracted from 2026-10-19 "New devices are data, not code" is the right bet for e-paper firmware and the whole-stack openness (SDK + firmware + KiCad + BOM) is rare. Want to see the SDK drive a controller family GoodDisplay didn't make, de-link get a license and resume commits, and whether CrossPoint's userbase follows the generalization or stays on the device-specific firmware.
superhq macOS-only alpha, three-week-old repo (2026-04-04), explicit "not production-ready," AGPL-3.0 single-vendor 2026-07-24 Auth gateway design is genuinely novel; want to see if it survives past the launch hype and broadens beyond Apple Silicon. shuru-sdk underneath is the more durable half.
gova Two-day-old repo (2026-04-22), single author, no tagged releases, pre-1.0 API, Fyne fallback on Windows/Linux 2026-07-24 Explicit-Scope reactivity for Go GUI is a genuinely novel ergonomic bet; want to see if Windows/Linux get native integrations and whether the API settles before committing.
stash One-day-old repo (2026-04-24), single author (established Go dev with 247 repos), three v0.x releases in first 24h, stages 4–9 of consolidation pipeline (causal/goals/failures/hypotheses) unproven in real use 2026-07-25 Nine-stage consolidation pipeline goes substantially further than hippo-memory / mempalace; want to see if the higher-order stages produce useful output or just thrash, and whether anyone besides the author ships a non-trivial integration.
atomicapp Five-month-old repo (2025-11-25), single author, SQLite-only datastore (not plain markdown on disk), 1.3k stars in 5 months, very active 2026-07-25 The LLM Wiki pattern shipped as one binary instead of bolted onto Obsidian. Want to see if Ken sustains the velocity, whether contributors arrive, and how data export / migration holds up since you can't grep your way out of a SQLite-only store.
lfk Six-week-old repo (2026-03-15), single author, pre-1.0 (v0.9.30), 320 stars, 30 releases in six weeks, no v1 commitment yet 2026-07-26 k9s alternative with a genuinely different navigation model (Miller columns + owner-based hierarchy) and built-in ArgoCD/Argo Workflows/Helm/KEDA/External Secrets integration. Want to see if release cadence holds, whether v1.0 ships, and whether a second contributor appears.
vera Ten-week-old repo (2026-02-22), single author (Alasdair Allan), pre-1.0 (v0.0.127, 127 releases), 261 stars, deliberately anti-ergonomic design (no variable names) 2026-07-29 Genuinely novel design bet β€” typed De Bruijn slots + mandatory contracts + Z3 verification, all aimed at LLM-as-author failure modes. Want to see if anyone besides Allan writes nontrivial Vera, if VeraBench gets independent reproductions, and whether the verified-MCP-server milestone lands.
goose-relay-vpn New repo by single author (kianmhz), CLI-only (no Android/iOS yet), Persian-speaking-market focus, AES wrapper unaudited 2026-07-29 Architecturally cleaner than masterhttprelayvpn (end-to-end AES, Google sees no plaintext) but trades the appeal "no infrastructure required" for a $4/month VPS. Want to see if independent contributors arrive, mobile clients ship, the AES integration gets reviewed, and whether the threat model holds against Iran's TSPU-equivalent.
bluetui One-day-old repo (2026-05-01), zero stars, single author, Russian-only README, no tagged releases, packaging is pip install -r requirements.txt + shell wrapper 2026-08-02 Real Bluetooth HID remote for Smart TVs (Profile1 with UUID 0x1124, raw L2CAP on PSM 17/19, Consumer Control + Keyboard descriptor) is a genuinely unusual implementation; rest of the BT manager is utility. Want to see if anyone else lands a PR, packaging tightens (PyPI), README gets translated, and TV compatibility expands beyond the author's TCL test.
openwarp Three-day-old fork (2026-04-29), 422 stars on momentum, single-vendor (zerx-lab, no other notable repos), no public release binaries, "early development" banner on the site, must cargo build --release from source 2026-08-02 Warp's first credible BYOP fork β€” the genai-based 6-native-protocol layer with reasoning passthrough is genuinely better than an OpenAI-compat shim, and AGPL/MIT mirroring upstream keeps it merge-able. Want to see if a binary release lands, whether zerx-lab actually keeps up with upstream Warp's release cadence, and whether a second contributor appears before the launch hype dissipates.
agent-skill-linter Three-month-old repo (2026-02-12), 2 stars, single author (William-Yeh), no editor integration / pre-commit hook, depends on the cross-harness "Agent Skills" registry (agentskills.io) catching on 2026-08-02 The deterministic-lint-for-skill-publishing instinct is right (same argument as impeccable's detector); 21-rule scope is well-chosen. Want to see if agentskills.io registry adoption picks up, if a second contributor lands, and whether the linter expands past Python skills (deeper rules currently Python-only).
acai Single-author org (~13β˜… across cli/server/docs), all repos created in last 2-3 months (CLI 2026-03-22, server 2026-02-27), CI/CD wiring still roadmap, hosted dashboard is the review surface with no published business model ("free forever maybe") 2026-08-03 The ACID convention (numbered acceptance criteria referenced from code) has legs independently of acai.sh, but the toolkit itself depends on the hosted dashboard for review. Want to see a notable third-party project adopting feature.yaml publicly, the CI hook examples landing, a second contributor in the org, and the hosted/self-hosted business model clarified.
surf Single-author project (1.7kβ˜…) tracking browser TLS/QUIC fingerprints; viability hinges on keeping pace with Chrome/Firefox releases and uTLS upstream; non-idiomatic g.Result API from enetx/g 2026-08-04 Bus-factor-1, browser-version treadmill (Chrome ships every ~4 weeks, uTLS lags), unusual API. Want to see if a second maintainer arrives, fingerprint coverage holds against current Chrome/Firefox, and whether the Result-type dependency stays optional.
tweakcc CC version-coupled binary patcher; whether Anthropic ships CC versions tweakcc can't keep up with, and whether Piebald keeps maintaining it 2026-08-04 Patches a closed third-party CLI β€” value evaporates the moment upstream changes outpace patches. Single-vendor maintenance bus factor. Want to see how the patch set tracks new Claude Code releases and whether the LIEF-based native-binary path holds.
oh-my-pi Single-maintainer fork (Can BΓΆlΓΌk) tracking upstream Pi (Mario Zechner's pi-mono); risk of bitrot if maintainer bandwidth drops 2026-08-04 Batteries-included counterpoint to Pi's minimalism. Want to verify sync cadence with upstream, whether the bundled tools (LSP, IPython, browser, SSH, multi-model role routing) drift from pi-mono, and whether a second contributor appears.
tilde-run Closed-source SaaS in private preview, no published pricing, single vendor; container-level (not microVM) isolation; homepage was carrying a prompt-injection payload at ingest 2026-08-05 Transactional commit/rollback for agent runs is a genuinely useful framing and the only product I've seen ship it as managed infrastructure. Want to see GA pricing, threat-model statement, whether self-host is on the roadmap, and whether the "agent-first RBAC" claim survives third-party scrutiny.
agent-skills-eval Four-day-old repo (2026-05-06), single author (darkrishabh), 257β˜… at ingest, MIT, TypeScript-only, costs scale 2Γ— per iteration 2026-08-10 The with_skill / without_skill baseline split is the right methodology for measuring skill lift, and spec compliance is honest. Want to see contributor count, whether judge-flakiness mitigation patterns emerge, integration with the rest of the agentskills.io ecosystem, and whether the agentskills.io spec stays stable enough for tools at this layer.
routing-run New service (no track record), license claimed open-source but repo not linked from homepage, per-request pricing only beats per-token at large context sizes, "zero logging" is a contractual claim 2026-08-10 The per-request pricing model and OSS-routing-infrastructure claim are differentiators if both hold. Want to see the actual repo, customer references, whether per-request economics survive OpenRouter's pricing evolution, and a privacy posture worth verifying.
re_gent Pre-v1.0 (POC-completeness per README), 265β˜… at ingest, Apache-2.0, Claude-Code-only today, planned features (rewind/fork/multi-tool) not built 2026-08-10 Per-prompt agent VCS solves a real problem named in agent-principal-agent-problem. Want to see v1.0, Cursor/Cline adapters, integration with PR review tooling, second contributor, and whether other coding-agent ecosystems ship comparable provenance primitives that obviate it.
floci Young project (org floci-io), Docker image just renamed hectorvent/floci β†’ floci/floci, riding the LocalStack-Community-sunset moment, several services still stubs (Bedrock, Textract), fidelity is best-effort 2026-08-10 Genuinely fills the gap LocalStack Community left in March 2026 β€” MIT, no auth token, broader service coverage, real-Docker-container fidelity for stateful services. Want to see contributor count beyond the founder, release cadence, whether the "drop-in LocalStack replacement" claim holds across real init scripts, and whether the stub services get filled in.
gosentry Day-of-announcement Go toolchain fork from a single security firm (trail-of-bits); must track upstream Go releases on an ongoing basis; LibAFL runner adds a Rust toolchain to the build env; bug-find list so far is L2/crypto-shaped 2026-08-12 "Same harness, stronger engine" is the right framing for a fuzzing fork β€” backward-compatible with testing.F, plugs Go's four worst gaps (path constraints, struct/grammar inputs, Go-specific bug classes, reporting). Want to see upstream-Go-tracking cadence, second contributor outside ToB, whether the LibAFL integration survives Go's quarterly toolchain churn, and whether anyone reproduces the bug finds on non-L2 codebases.
mnemonik Young project (single-vendor mnemonik-xyz), Solana + Arweave hard dependency in full mode, embedding choice commodity (TurboQuant + default ONNX), monetisation surface unclear past free local mode 2026-08-13 Verifiable-memory framing fills a gap the recall-side memory tools (hippo-memory / mempalace / stash) explicitly don't address. Want to see if anyone outside the team ships an integration, whether the COSE/blake3/Ed25519 + Solana anchor flow holds up under audit, and whether the embedded recall quality keeps up with dedicated retrieval tools.
statewright New repo (88β˜…), single-team launch, Rust engine open-source under Apache-2.0/FSL but managed cloud is the review surface, advisory enforcement on Cursor only, research validation is a 5-task SWE-bench subset not the full 2294 2026-08-13 State-machine guardrails for agents are a clean structural answer to flailing-on-too-many-tools failure modes β€” the 2/10 β†’ 10/10 SWE-bench-subset numbers on local models are the strongest evidence in the wiki for tool-space constraint as the lever. Want to see contributor count, whether the engine sees non-cloud adoption, full SWE-bench reproduction, and whether the managed-cloud / self-hosted balance settles.
needle 26M-param model from single team (Cactus Compute), four days old at ingest, single-shot function-calling only, English-only by training, finetuning is the recommended path 2026-08-13 The Simple Attention Network shape (no-FFN encoder, ZCRMSNorm, gated residuals, shared embed/output) is a real research contribution that ships with weights and a finetuning pipeline. Want to see if anyone outside Cactus ships finetuned checkpoints, multilingual or multi-turn variants land, and whether the 6000/1200 toks/sec numbers transfer off the Cactus runtime.
crofai Single-domain hosted LLM provider with no observable track record, closed-source backend, Q4_0 baseline quantization on production models, public pricing page is the only documented surface 2026-08-13 Quantization disclosure as a column is genuinely useful and the prices are low. Want to see third-party usage reports, whether quality complaints surface on the aggressive Q4_0 baselines, sustained catalog updates, and whether the "cheapest" claim holds against repricing rounds from DeepInfra / Fireworks / Together.
forgezero Two-month-old repo (2026-03-15), single author (alexvoste), 5 stars, macOS "in progress" and Windows "experimental," aggressive 1.5β†’1.9 release cadence in two months suggests pre-adoption iteration 2026-08-19 Strict-warnings-and-sanitizers-by-default + pre-link symbol check is a genuinely opinionated bet on a niche (assembly + bare-metal C) most modern build tooling ignores. Want to see a second contributor, macOS/Windows parity, whether the no-escape-hatch strict defaults survive real user pushback, and whether the project finds an audience beyond hobbyist OS-kernel writers.
pve-microvm v0.3.12, single author (Rui Carmo), ~320 stars, Proxmox-only, structurally coupled to patching upstream qemu-server Perl internals it doesn't control (every PVE upgrade can break it) 2026-10-18 Genuinely useful "container-speed VMs with actual isolation" on a homelab node, and the host-provided-kernel / OCI-rootfs design is clean. Want to see whether the version matures past 0.3.x, a second maintainer appears, the QEMU-10.x MMIO-for-Linux and GPU/vIOMMU-passthrough gaps get contributor patches (Carmo is explicitly asking), and whether any upstreaming to Proxmox happens.
lazypi v0.6.3, single author (Rob Zolkos), ~373 stars, MIT, value is pure curation over a fast-moving third-party Pi package ecosystem it doesn't control 2026-10-18 The LazyVim-for-Pi installer (vs oh-my-pi's fork) is a clean answer to "vanilla Pi ships empty," and it's the best evidence Pi's community-package bet worked. Want to see whether it gains a second maintainer, whether the curated catalog stays fresh as packages come and go / get abandoned, and whether Pi's ecosystem keeps growing enough to justify a curation layer at all.
xslang Single-author project at v1.2.32, limited public history; ambitious cross-target claims (Linux/macOS/Windows/WASI/iOS/Android/ESP32) need real-world validation; site doesn't make license obvious 2026-08-21 Single 2.9 MB binary containing the whole toolchain (JIT + LSP + debugger + REPL + transpilers) across that many targets is unusual together. Want to see external adoption (any non-author projects), iOS/Android/ESP32 claims under real scrutiny, a clearly stated license on the site, and whether the JIT/VM benchmark numbers hold against fresh comparisons.
freenet From-scratch Rust rewrite of Freenet (formerly Locutus); production adoption signal is the demo network, not dApps with users; small team funded by grants/donations 2026-08-22 Small-world-ring + browser-served decentralized apps is a structurally different bet from Tor / Yggdrasil / IPFS, and the new Freenet has shipped a working network plus developer manual. Want to see any non-demo dApp with users, contributor count beyond core team, and stable funding past the current grant footing.
sandbox-agent-sdk Built on Docker's undocumented sandboxd.sock API which can break in any minor release; macOS/Windows-only because Docker Sandboxes is; production signal not yet visible 2026-08-22 Lets you escape Docker Sandbox's six-agent whitelist and run arbitrary containers behind microVM isolation β€” the actual unlock people will want. Want to see whether the upstream API stays stable, second contributor, integration patterns beyond the demo agents, and whether Docker formalizes the API surface.
kata Early public preview, no tagged release yet, command contracts explicitly unstable, 210β˜…, no auth in local or remote mode, import always recreates the DB 2026-08-22 The keep-issues-outside-the-repo bet (vs. beads' Dolt-in-repo) is a clean alternative for non-git or clean-footprint workflows. Want to see contracts stabilize into a tagged release, a second contributor, and whether the planned shared-server mode ships with actual auth. Part of the kenn-software-suite.
msgvault Self-described alpha (v0.14.x), README warns storage format and CLI flags may change without notice; 1.8kβ˜…, single-vendor (Kenn Software) 2026-08-22 The SQLite-for-store + DuckDB-for-analytics split is the right shape for lifetime-email archival and the vector/hybrid search is genuinely useful. Want to see the storage format freeze, the alpha label drop, contributor count beyond the vendor, and whether incremental Gmail History sync holds up over very large mailboxes. Part of the kenn-software-suite.
mindwalk Cross-session cumulative history; adapters beyond Claude Code/Codex; whether the spatial view beats a text diff 2026-10-16 Single-author, brand-new (July 2026), value proposition still contested on HN
zerofs WAL / multi-writer support; release cadence; second contributor 2026-10-16 Single-author young storage project; single-writer, no WAL yet; author-written benchmarks
reaction External security audit + sustained activity 2026-10-16 Single maintainer (ppom), no audit, Linux-only, young v2 Rust rewrite; runs as root wiring firewall rules
gap fred integration / graduation from jam proof-of-concept; macOS support 2026-10-16 Experimental single-author jam project; the diff engine is meant to live inside fred, not ship standalone
commonforms Second maintainer + a clear/standard license; dataset-prep code leaving WIP 2026-10-16 Single-author repo (Joe Barrow), depends on the FFDNet models, custom NOASSERTION license
moonstone 1.0 release, Windows support, real-world adoption vs Lux/Nix 2026-10-16 Young single-author Zig project, POSIX-only, pre-1.0 with shifting APIs
flint-chart API stabilizing; Python package release; token/correctness benchmarks 2026-10-16 Early Microsoft Research release, evolving JSON spec, unreleased Python port, no benchmarks yet
kastor Second codegen target + first real hosted-platform provider beyond in-memory 2026-10-16 Early proof of concept, single author, 57 stars; only LangGraph codegen + in-memory target work today
sp4rk First tagged release + adoption beyond the author (0 stars, README says early alpha) 2026-10-16 Pre-release, single-author (v0lka), APIs may change without notice
showagent Star traction, release cadence, second maintainer 2026-10-16 Single-author, early (34 stars); promising cross-agent session conversion but unproven
dockerscan v2.1 offensive/registry features shipping; license terms staying usable 2026-10-16 Single-author; proprietary source-available license (not OSS), commercial-competitive use restricted
tc-lang ~40 stars, single author (Alonso VM), pre-1.0, MIT README badge but no SPDX license file detected by GitHub, "make an OS with it" aspiration 2026-10-19 Transpile-to-C11 minimalist systems language with genuinely unusual ergonomics (strun, fat pointers, host/DLL hot-reload, async channels). Want to see a second contributor, whether stage1 self-hosting happens, the license clarified, and any users beyond the author.
kanbots New product, open-core, no public source repo (landing page only), desktop-app maturity unproven, single vendor 2026-10-19 Worktree-per-card parallel-agent kanban with cost caps and decision-gating is a clean framing. Want to see adoption, whether the OSS build really ships all features, contributor count, and whether the open-core split holds.
claude-code-workflow-creator 85 stars, single author (Ray Amjad), no LICENSE file, depends entirely on Claude Code's unreleased env-gated Workflow tool (CLAUDE_CODE_WORKFLOWS=1) 2026-10-19 The deterministic-JS-orchestration model matches how the unreleased Workflow tool actually works. Want to see whether Anthropic ships the tool publicly, whether the skill tracks the released API, a license appearing, and a second contributor.
oak No stated license, binaries only for macOS-arm64 and Linux-x86_64, and a hosted server-side main in the default workflow 2026-10-21 Snapshotting large repos without a full clone, branch-per-task, and mounting rather than cloning is the right shape for agents that check out a tree per task. Want to see a license, platform breadth, and whether the hosted main is optional in practice.
activegraph Single-author research project shipped alongside its paper; compaction, concurrency and any measured task result are all listed as future work; replay is linear in log length 2026-10-21 Event-sourced log as source of truth with the graph as a deterministic projection buys replay, cheap forking and per-object provenance that retrieval-based memory cannot. Want to see checkpointing land, a concurrency story, and any empirical result beyond the worked example. See log-is-the-agent.
bramble Repo public 2026-06-01, 249β˜…, single vendor, no external audit of the crypto core, README discloses LLM co-authorship on security-critical code, scheduled cloud backups desktop-only 2026-10-21 LUKS-style key slots and envelope encryption in one Rust core shared by extension and both mobile apps is a clean vault design, and passkeys-as-entries with peer-to-peer sync is what hosted managers can't offer. Want to see an independent audit, a second contributor, cloud backups on mobile, and the vault format freeze before it holds real credentials.
quality-md Six weeks old (2026-06-11), 29β˜…, single vendor, no third-party implementation of the spec, evaluation is LLM judgement against a rubric 2026-10-21 Declaring the quality rubric in-repo sits at the right level above per-feature acceptance criteria (acceptance-criteria-ids, acai), and the numbered .quality/evaluations/ history is a good shape. Want to see adoption by projects unrelated to the authors, a second implementation, and whether the reports beat an organized code review.
delirehberi-news 16β˜…, single author, last push 2026-07-03, no releases, entire stack is Cloudflare-proprietary (Workers, D1, Vectorize, Workers AI), sources hardcoded to HN/Lobsters/Reddit 2026-10-21 An embedding-ranked personal feed keyed to a Nostr pubkey rather than an account is worth copying even if this repo goes quiet. Want to see whether the author keeps it alive, whether anyone else deploys it, and whether the Cloudflare coupling loosens.
bento No public source repository and no stated license; the advertised collaboration is absent from the demo; the file self-updates from a signed manifest by default 2026-10-21 Deck, viewer and editor in one self-saving HTML file with the runtime pinned inside is a real answer to document rot, and live charts that morph between forms are more than a gimmick. Want to see the source published, a license, and what collaboration actually requires on the network.
allyourcodebase Every package is pinned to the latest tagged Zig while the build-system API is still unstable (returning-to-zig records 0.17 being expected to break every project's build), and per-repo maintainership is asked of volunteers with nothing enforcing it β€” two moving targets per package, no visible signal when one stops being tracked 2026-10-27 The pristine-tarball + build.zig strategy and the "upstreaming must not add system dependencies" archive condition are worth stealing regardless. Want to see how the org fares across a breaking Zig release, whether package count keeps growing or the long tail bit-rots, and whether any upstream has actually adopted and let a repo be archived.
audio-cassette-simulation Whether the per-tape directory duplication collapses into one script taking a profile argument, whether filter parameters get documented outside the scripts, and whether more tape types land 2026-10-27 Single author, no releases, README-only documentation, and the profiles are described rather than measured against real decks. The six named tapes (including Soviet MK-60) are the interesting part; the packaging is not.
cursor-bridge Whether Cursor tolerates third-party clients on Auto quota, and whether the Anthropic-API translation survives either vendor's changes 2026-10-27 Lifespan is set by another company's business decision rather than its own maintenance. Single author, no sandboxing, single account at risk. Worth noting you get Cursor's Auto model β€” "keep the client, change the model", not free Opus.
deltafin A native CUDA MXFP4 MoE kernel (routed experts still fall back to the CPU kernel on NVIDIA) and the promised quality harness measuring average NLL against the official API 2026-10-27 Research artifact by the maintainer's own description. Every headline number comes from one M1 Max, and the only cross-platform data point is a single unreplicated community run. Streaming experts from disk is the idea worth tracking whatever happens to this implementation.
feynobg Independent reproduction of the eight-benchmark ranking, the withheld throughput/latency/memory numbers, and whether the model line outlives the launch 2026-10-27 87β˜…, one startup, SOTA claim self-published with only the UHRSD-TE pair (0.981 vs BiRefNet 0.957) spelled out. NoBg as a common training/serving interface may prove more durable than the model.
hartos A green run of the 19 nixosTest VM checks, any VERIFICATION.md row settled on real hardware, and one payment settling through collective_earning.py 2026-10-27 Public alpha pitched as an OS with boot untested, learning core closed and compiled in a private repo, no revenue path exercised, no PyPI package, Python 3.10/3.11 only. The README discloses nearly every gap itself, which is candour worth reading as part of the pitch β€” judge it by the unrun tests.
pgsimcity The self-assessment gap closing β€” the live site says "early, unreviewed prototype" while the README claims three specialist review rounds and a 234-test suite β€” plus a 1.0 or stable accuracy statement, and touch controls verified on real devices 2026-10-27 Single-author 0.x educational model of PostgreSQL that says outright it "almost certainly contains inaccuracies", so its value as a reference rests entirely on the review claims holding.
yap Whether it stays viable across macOS releases, and whether a cross-platform equivalent appears 2026-10-27 Hard floor of macOS 26 Tahoe on Apple Silicon, Xcode 26 to build, entirely dependent on Apple's SpeechAnalyzer API β€” a company side project with a deliberately empty roadmap. The no-bundled-model bet is the thing to watch; it either ages very well or breaks on one OS release.

| boffin | The routing engine staying a black box β€” ParselFire Core decides which constraints reach which edit and the README never describes it β€” plus an A/B comparison rather than three hand-picked case studies, and a second contributor | 2026-10-27 | 37β˜…, single author, first commit June 2026. Routing architectural constraints per-edit instead of dumping them in a prompt is the right instinct; whether it works is currently unevidenced. | | mtproxy-reanimation | Whether Telemt keeps absorbing its features β€” 3.4.18+ already overlaps and the README tells you to pick one β€” and whether any claimed figure gets measured independently | 2026-10-27 | Repo created 2026-06-10, one author's bash, 294β˜…, curl \| sudo bash install writing systemd units and kernel firewall rules, Docker bridge and zapret2 untested per its own README. Censorship tooling has a short half-life; the zapret2 wscale constraint is the durable idea. | | onecli | A published threat model and a security review, plus identity providers beyond Google for multi-user | 2026-10-27 | 2,924β˜… argues against watchlisting on traction, but it terminates TLS and holds every secret with nothing published about how. Hiding a key also does not reduce what the key can do β€” see credential-compartmentalization. | | palmier-pro | Whether the MCP tool surface gets documented and whether the OS floor relaxes | 2026-10-27 | macOS 26 on Apple Silicon only, generative core closed and subscription-gated. Driving a desktop GUI's document model over MCP is a shape the vault has little of, which is why it is worth tracking at all. | | petals | A binary question: does upstream resume or does the repo get archived? Last commit 2024-08-25, last release v2.2.0 (2023-09-06), 113 open issues, not archived. Also whether health.petals.dev shows any live hosts | 2026-10-27 | Still the clearest design for volunteer-computing LLM inference, and a private swarm across owned machines survives the trust objection even if the public one is dead. The landing page presents a 2023–2024 model lineup in the present tense. | | pullrun | Independent reproduction of the benchmarks, a second contributor, and the CRI shim leaving beta | 2026-10-27 | Repo created 2026-06-11, 114β˜…, 1 fork, single apparent team shipping 11 Rust crates plus 4 Golang components; every number and "only runtime that does X" claim is self-reported, and a CLA sits on an Apache-2.0 project. The zero-copy DAG store with refcounted cascade-delete is worth stealing regardless. | | pvz-ps2 | Whether it gains a release and any evidence of running on real hardware rather than only PCSX2 | 2026-10-27 | Six commits over two days (2026-07-09/10), no release, no screenshot, README inherited from upstream with a platform table that doesn't list PS2. The code is real and specific β€” hand-written GL subset over gsKit, VRAM allocator, PSMCT16/T8 conversion β€” the outcome is unrecorded. | | remux | An App Store release or tagged builds, whether the prebuilt GhosttyKit dependency keeps tracking upstream Ghostty, and whether anything beyond iPhone is planned | 2026-10-27 | 212β˜…, two contributors, repo created 2026-04-21, TestFlight public beta with zero GitHub releases. File preview and localhost preview are genuinely new for a mobile terminal, and direct SSH with no relay and no account is the right trust model. | | wanix | A 1.0 β€” it is pinned at 0.4.0-rc2 β€” and evidence of use outside the author's own demos; workbench extensions and live editing are still documented as future work | 2026-10-27 | Single-shop project (tractordev) since 2023-10, 797β˜…, steady rather than fast. Per-process namespaces are a better sandboxing primitive for untrusted browser-side code than ad-hoc WASI capability plumbing, and this is the only implementation aimed at the web. | | whetuu | A second contributor, a 1.0, whether zero-config survives feature requests, and whether the 39-toolchain detection list keeps pace. No Windows builds | 2026-10-27 | Three-week-old repo (2026-07-05), single contributor, v0.1.10, 35β˜…. Careful in the places these tools usually skip β€” --no-optional-locks so it never contends for index.lock, bounded probes, control bytes defanged, history store forced to mode 600. The no-config stance is both the differentiator and the risk. |

| kata | Command contracts settling β€” the README calls them explicitly unstable β€” and the first tagged release; both were the stated reason for tracking it | 2026-08-20 | Early public preview where the CLI, daemon and TUI all work, so the gate is interface stability rather than completeness. Keeping issue state outside the repo is the design bet worth watching. Row reconstructed 2026-07-29: the page was tagged and listed from 2026-05-22 but never had one. | | msgvault | The storage format stabilizing and the alpha label dropping β€” v0.14.x still ships a format-may-change warning | 2026-08-20 | An archive you cannot migrate out of is worse than no archive, so the format warning is the whole gate here. SQLite FTS5 plus DuckDB/Parquet over the same corpus is the part worth having. Row reconstructed 2026-07-29: the page was tagged and listed from 2026-05-22 but never had one. |

| agent-manager | Status detection keeping pace with eight fast-changing agent CLIs, a second maintainer, and the cost tracking the README lists as missing | 2026-12-12 | Repo created 2026-07-15, single author, 432β˜…, heavy launch promotion (Trendshift, Product Hunt, Peerlist). Line-comment diff review sent back to the agent is the part worth keeping. | | engrim | Evidence beyond the author's own 105-session case study, a second contributor, and some review step for agent-written records | 2026-12-12 | Created 2026-06-19, single author, 241β˜…. The "standard" has no spec, the savings figure has no baseline, and setup edits every agent's global config. | | pigeon | Activity after launch week, a PyPI release, a security review of the pass format, and any adoption | 2026-12-12 | One week old β€” every commit is from 2026-09-06 β€” 41β˜…. A security primitive with no review yet, and key handling is not shown in the README. | | cloud-in-a-bottle | Catalog growth, documented inter-app permission APIs, and the self-hosted path staying first-class beside the managed tier | 2026-12-12 | Launched publicly 2026-09-05. Single corporate sponsor (Imbue) whose revenue is the managed product; the app catalog is small by its own account. | | vps-audit | Checks extending beyond Debian/Ubuntu, and thresholds that are less arbitrary | 2026-12-13 | One distro family; every PASS/WARN/FAIL verdict comes from hand-picked thresholds. | | darling | GUI support β€” whether AppKit and Metal-over-Vulkan run non-trivial GUI apps β€” and a clear statement on CPU architectures | 2026-12-13 | GUI support is "basic experimental" for simple apps by the project's own account. | | tare | A second maintainer, tagged releases, and whether it survives the next Claude Code log-format change | 2026-12-13 | One contributor, created 2026-08-12, parses Claude Code's undocumented session logs; the owner Kelviq sells usage billing. |

Graduated

(none yet)

Abandoned

(none yet)