Validation strategy
Status: proposed research plan. No experimental results are recorded here.
Question and comparison
Test whether an agent with XELON completes tasks better than the same agent alone, and whether any benefit justifies additional time and compute.
Use a documented task set with explicit acceptance criteria. Coding tasks are a possible starting point; an authentication-system task could require tests, security checks, interface compatibility, and a successful build. Automated checks do not establish comprehensive security.
Compare:
- Agent Alone: the selected agent using its documented normal workflow.
- Agent + XELON: the same agent with the proposed evaluation, learning, optimization, and stopping cycle.
Keep agent/model versions, task inputs, tools, environment, and permissions comparable. Record the agent’s existing retry or planning behavior so that the baseline is not artificially weakened. Include an equal-budget comparison and, where useful, a single-attempt baseline. Separate extra-compute effects from protocol effects.
Measures
| Measure | Proposed reporting |
|---|---|
| Success rate | Fraction satisfying all required acceptance criteria, with uncertainty |
| Quality | Predeclared rubric and per-criterion outcomes, with reviewer agreement where applicable |
| Time | End-to-end elapsed time, including evaluation and orchestration |
| Token/compute cost | Input/output tokens, calls, relevant compute usage, and monetary cost where known |
| Failures | Failure categories, regressions, exhausted budgets, and stopping reasons |
Report absolute outcomes and differences, distributions, and sample sizes. Retain failed runs and missing measurements. Distinguish measured usage from cost estimates, and record pricing assumptions if monetary estimates are used.
Experimental discipline
- Choose tasks, criteria, budgets, stopping policies, and the analysis plan before evaluation.
- Separate development tasks from held-out evaluation tasks to limit tuning leakage.
- Run paired comparisons with repeated trials where agent randomness matters; randomize execution order and reset workspaces between runs.
- Use independent evaluators where feasible. Keep held-out checks outside the agent’s control and blind human reviewers to the condition where practical.
- Preserve configuration, versions, seeds where supported, artifacts, evaluator outputs, usage records, and deviations from the plan.
- Analyze uncertainty, failure modes, and cost-quality tradeoffs; publish the limits of the evidence along with any future results.
Sample size and thresholds must be justified before making improvement claims. A favorable illustrative score is not a measured outcome.
What caused a change?
Ablations should compare retry-only behavior, evaluation with history, structured learning, and strategy changes under comparable budgets. This can help distinguish benefits from additional attempts, better evaluation, or actual adaptation. Inspect whether the next strategy used the recorded learning and whether that change improved the outcome; logs alone cannot establish causality.
Decision criteria and limits
Before testing, define the minimum practically useful success or quality gain, acceptable time/cost overhead, and unacceptable regressions for the chosen domain. No numeric improvement target is established by this founding package.
Negative and inconclusive findings are valid outcomes. Risks include benchmark contamination, evaluator bias, small samples, nondeterminism, reward gaming, unequal budgets, and poor transfer across agents or domains. Results from one agent or task set must not be generalized without further evidence.