Recursive Self-Improvement (RSI) AI Research Automation as a State-Space System Hypothesis Generation, Code Generation, Experiment, Evaluation, Selection, and Recursive Self-Improvement
Updated 2026-10-04 · NEXMASON ANITEX▶ Open interactive ANITEX · equations and animations
This accessible text edition preserves the document's narrative. See the interactive edition for typeset equations, diagrams and playback.
State-Space Model of an Autonomous AI Research Loop
An autonomous AI research process can be formulated as a discrete-time nonlinear state-space system.
The research loop is
Hypothesis Generation Code Generation Experiment Evaluation Selection State Update
The essential property is that the result of one research cycle changes the internal state from which the next research cycle begins.
Research State Vector
Let the complete research state at iteration k be
x_k = K_k H_k C_k X_k E_k B_k A_k ,
where
K_k = accumulated scientific knowledge, H_k = hypothesis state, C_k = code and implementation state, X_k = experimental observations, E_k = evaluation state, B_k = belief state, and A_k = research agenda and priorities.
Thus,
x_k = [K_k,H_k,C_k,X_k,E_k,B_k,A_k]^T
represents the persistent memory of the autonomous researcher.
General State-Space Representation
The research process can be written as
x_k+1 = f(x_k,u_k,w_k)
with observation equation
y_k = g(x_k,u_k) + v_k.
Here,
u_k = research action, w_k = process uncertainty, and v_k = measurement uncertainty.
The observable vector may contain
y_k = accuracy loss runtime memory robustness generalization safety score .
This distinction is important because the true quality of a research state is generally not directly observable.
Stage I: Hypothesis Generation
Given the current research state x_k, the AI generates candidate hypotheses
h_k^(i) _H(hx_k), i=1,,N.
Thus,
H_k = \ h_k^(1), h_k^(2), , h_k^(N) \.
Each hypothesis can represent a proposed change in architecture, algorithm, optimizer, prompt, memory, tool policy, or even the research process itself.
A hypothesis utility function may be defined as
V_H(h_i) = N_i + P_i + I_i - C_i,
where
N_i = novelty, P_i = estimated plausibility, I_i = expected information gain, C_i = experimental cost.
The hypothesis selected for implementation is
h_k^* = _h_iH_k V_H(h_i) .
Stage II: Code Generation
The selected hypothesis is transformed into an executable implementation.
Let
c_k = _C(h_k^*,x_k).
The implementation operator therefore performs
h_k^* _C c_k .
More generally,
c_k = G_C(h_k^*,K_k,C_k),
where C_k contains the existing source code and K_k contains relevant technical knowledge.
The generated code must satisfy validity constraints
_j(c_k)=1, j=1,,m,
such as compilation, unit tests, interface compatibility, and safety constraints.
Only valid implementations proceed to experimentation.
Stage III: Experiment
The generated implementation is applied to the experimental environment.
Define
e_k = E(c_k,d_k,r_k),
where
d_k = experimental data, and r_k = allocated computational resources.
The experiment produces observations
z_k = z_1,k z_2,k z_m,k .
For example,
z_k = benchmark accuracy validation loss execution time memory consumption generalization score .
The experimental observation model becomes
z_k = g_E(x_k,c_k) + _k
where _k represents experimental noise.
Stage IV: Evaluation
The experimental observations are converted into an evaluation score.
Let
J_k = J (z_k,z_baseline).
A multi-objective evaluation may be written as
J_k = w_1 P_k - w_2 C_k + w_3G_k + w_4R_k - w_5S_k,
where
P_k = performance improvement, C_k = additional computational cost, G_k = generalization, R_k = robustness, S_k = safety risk.
Therefore,
J_k = w^Tq_k
for a quality vector q_k.
Stage V: Selection
Suppose several candidate experiments have been evaluated:
\ J_k^(1), J_k^(2), , J_k^(N) \.
The preferred candidate is
i^* = _i J_k^(i) .
Selection should also require an acceptance threshold.
Define
a_k = 1, & J_k^(i^*) > J_baseline + , 0, & otherwise.
Here is a minimum improvement threshold.
If a_k=1, the modification is accepted. If a_k=0, the previous system is retained.
Thus,
_k+1 = a_k_k^* + (1-a_k)_k.
This produces a validation-gated research loop.
Belief-State Update
Experimental evidence must modify the system's beliefs.
Let
B_k()
represent the current belief distribution over research hypotheses.
After observing z_k, Bayesian updating gives
p(z_1:k) = p(z_k) p(z_1:k-1) p(z_kz_1:k-1) .
Thus negative experiments are also informative.
A failed experiment changes
B_k B_k+1
and therefore affects future hypothesis generation.
Complete Research Transition Operator
The entire research cycle can be represented by the composition
F_R = F_U F_S F_E F_X F_C F_H
where
F_H = hypothesis generation, F_C = code generation, F_X = experiment execution, F_E = evaluation, F_S = selection, and F_U = research-state update.
Therefore,
x_k+1 = F_R(x_k) .
Expanded,
x_k+1 = F_U ] ] ] ] .
This is the fundamental state-transition equation of the autonomous AI research loop.
Recursive Self-Improvement Extension
Ordinary automated research modifies a target system .
Recursive self-improvement appears when the research system itself becomes part of the state being optimized.
Define
x_k = _k _k K_k B_k ,
where
_k = target AI parameters, and _k = parameters of the AI research process.
The transition becomes
_k+1 _k+1 = F(_k,_k,z_k).
The crucial RSI condition is
_k+1_k
because the mechanism performing future research has itself changed.
Therefore,
AI_k Research_k AI_k+1 Research_k+1 AI_k+2.
If the research capability is represented by
Q_k=Q(_k),
successful recursive improvement requires
Q(_k+1) > Q(_k) .
The AI has then improved not merely its task performance, but its capacity to discover subsequent improvements.
Closed-Loop Control Interpretation
The autonomous research system can also be interpreted as a feedback controller.
The desired research objective is
r_k.
The observed performance is
y_k.
The research error is
e_k = r_k-y_k .
The AI researcher acts as a controller
u_k = (x_k,e_k).
The research environment acts as the plant
x_k+1 = f(x_k,u_k).
The evaluator provides feedback
y_k = g(x_k).
Hence,
AI Research = Closed-Loop Adaptive Control
with the unusual property that the controller may eventually modify its own control law:
_k _k+1.
This produces a meta-adaptive control system.
Conclusion
The autonomous AI research loop can be modeled as a nonlinear, partially observed, closed-loop dynamical system.
Its fundamental state equation is
x_k+1=F_R(x_k)
where F_R contains hypothesis generation, code generation, experimentation, evaluation, selection, and state updating.
Recursive Self-Improvement emerges when the research mechanism itself is included in the optimized state:
(_k,_k) (_k+1,_k+1)
with
Q(_k+1)>Q(_k).
This transforms automated AI research into a meta-adaptive system in which the mechanism that discovers improvements can itself become the subject of subsequent improvement.