Not released yet. In active development and daily use.

Research Agenda: Verified Outcomes with Minimal Effective Teams

Can a small team solve harder problems with less time and compute? We plan to test that question against strong single-agent baselines.

Status as of 2026-09-26: Documented and deferred research agenda. Reliability, secure operation, and operator controls come first. The experiments described here have not started, and we claim no novel scientific result.

Objective and Falsifiable Hypothesis

The central research question is whether a small operator can solve previously unsolved problems, or achieve independently verified outcomes using materially fewer resources than a credible reference approach, by coordinating the smallest effective team of autonomous agents.

Falsifiable Hypothesis: Complementary specialization, persistent shared evidence, selective model escalation, and deterministic review gates may improve verified outcomes relative to total compute expenditure. Conversely, uncoordinated scaling introduces duplicate tool calls, correlated model hallucinations, compounding errors, and runaway spend. The research methodology explicitly accommodates the null hypothesis: where a single strong reasoning model or a minimal pair outperforms an adaptive multi-worker fleet.

The role of echo.cc in research

echo.cc supplies messages, task leases, and advisory file reservations. A research harness would need to add experiment bookkeeping, complete cost accounting, independent verification, and retention of unsuccessful attempts. Those research capabilities are proposed designs, not shipped guarantees.

Measurement & Resource Accounting Protocol

Rather than collapsing complex multi-agent behavior into a single opaque score, the proposed experiments would measure eight dimensions:

Research Measurement Protocol
DimensionRequired Measurement & Interpretation
Verified QualityIndependent checker pass rates, formal proof confirmations, and domain-expert review across repeated trials.
End-to-End Wall TimeComplete duration from campaign creation to final verification, including queueing, review cycles, and repairs.
Compute & TokensPrompt, cached, and generated token volumes; local GPU/CPU execution hours where measured. Unknown spend is never treated as free.
Comprehensive CostTotal financial accounting including coordinators, reviewers, retries, discarded exploration branches, and tooling infrastructure.
Human EffortSetup time, prompt adjustments, operational interventions, debugging, and post-run evaluation time.
Effective Team SizePeak concurrent inference requests, active worker sessions, and unique participants reported separately. Roster presence is not concurrency.
Coordination OverheadTotal messages and bytes transferred, idle worker capacity, context reconstruction delay, and dispatch backpressure latency.
Reliability & SafetyDuplicate tool side effects, interruption recovery rates, lease contention frequency, and accepted result lineage.

Empirical Status vs. Proposed Experiments

We maintain strict transparency between operational system demonstrations and future scientific research: