Build companies that learn.

RSIX turns every business action into learning, and every economic outcome into a better policy.

action outcome evaluation better policy repeat

Starting with AI-native commerce. Building toward Company Agent.

AI can act. Companies still can’t learn.

Ads, recommendation, pricing, creative, inventory — a company fires thousands of actions a day. The experience between action and outcome scatters into dashboards, SOPs, meetings, and the heads of whoever was in the room. Wins can’t be attributed. Failures repeat. Nothing feeds the next decision.

context hypothesis action outcome evaluation policy update software · records workflow agents · execute copilots · suggest ad platforms · one edge the loop closes RSIX · owns the loop end to end decision rights · experiment rights · outcome data · policy-update rights
Figure 1 — Every existing category sells a fragment of the loop: software records it, copilots suggest into it, workflow agents execute a step, ad platforms optimize one edge. The amber nodes — evaluation against realized profit, and the authority to update the policy — are sold by no one. They can only be owned.

The company acts every day. But the company itself does not get smarter.

Company Agent is the end state.
RSI is the engine.
Commerce is the training ground.

We are building the Company Agent — a company that runs itself and improves itself. But it must first have a real world: somewhere to act, to be wrong, to earn, to learn. Commerce is the first world we chose.

This does not require a frontier model that rewrites its own weights. Frontier models are the CPU — rented, replaceable, improving on their own schedule. RSIX builds the learning system around them: policy, memory, workflows, search strategy, experiment design, the evaluator, the agent graph.

L2 · verification economic evaluator · experiment ledger · safety, audit, rollback results L1 · agentic operations user agent · product agent · acquisition agent · commerce agents outcomes L0 · owned economic environment users · products · budgets · transactions · fulfillment · profit L3 · meta-optimizer improves: hypotheses experiment design search strategy memory & skills agent workflows the agent graph the evaluator itself verified evidence rewrites rebuilds
Figure 2 — The model sits inside L1; the loop is everything around it. Evidence flows up, improvement flows back down — into the agents, the experiments, and the evaluator that judges its own work. We optimize the loop, not just the model.

The model is the CPU. We build the learning system.

Company Agent is not a slogan. It is a falsifiable test.

In the same economic environment, can the system find better policy faster than the best human operators? That is the whole question. We run it as an experiment.

Elite human team VS RSIX

same budget · same traffic · same products · same constraints

Measured on one number

Economic Learning Velocity

attributable, sustained economic uplift — per dollar, per unit of time

Commerce is not our market. It is the training environment.

To train a company you need an economy that answers back: high-frequency actions, unambiguous outcomes, real-money rewards, a vast search space, experiments that settle in days, transactions that run end to end. Commerce is all six at once.

It is the logic that took reinforcement learning through Atari and Go — a fast, measurable world first. Except the reward is not a game score. It is profit, retention, returns, repeat purchase.

We begin in narrow, fast-feedback categories where the loop closes within days, then move into high-context purchases where a wrong buy is expensive. Models can be rented; traffic can be bought; the loop cannot be outsourced.

The reward is not a game score. It is profit.

The end state is a company creation engine.

Once the loop holds in one business, it replicates. A founder supplies four inputs; the system runs research, creation, acquisition, transaction, evaluation and improvement — and every new business unit starts from the accumulated policy of all the units before it.

a founder supplies goal · capital authority · risk boundary RSI core one loop, replicated verified experience compounds business unit 01 business unit 02 business unit 03 business unit N
Figure 3 — The dashed box is everything a person still supplies. Each new unit inherits the accumulated policy of every unit before it — less time, less capital, fewer people, better odds.

Soon every company will run on the same rented intelligence. When intelligence is abundant, it stops being the advantage. The advantage moves to learning speed — how fast real-world feedback becomes better behavior. Distribution, brand, supply chain, capital, talent: the old variables remain. AI-native companies add one more, and it compounds.

OpenAI trains models. RSIX trains companies.

Built by the people who ran these systems at scale.

The team behind RSIX trained the recommendation, advertising and generative systems of some of the largest consumer platforms on earth — and carried the P&L for what those systems earned. Closing an economic learning loop takes six disciplines that almost never share a room. This room has all six.

Frontier models×Recommendation×Advertising×Growth×Commerce×Supply chain

We are hiring the people who will close the loop.

[email protected]

Research and engineering: agent learning, evaluation, causal inference, recommendation, infrastructure.

[email protected]

Partnership, supply and capital.