Joint Reference Architecture · Wand AI × StorONE

Maximizing sovereign AI ROI through storage utilization.

Wand's sovereign backend raises the utilization of national compute. StorONE's Real-Time Tiering raises the utilization of the storage underneath it, serving the same AI workloads on a fraction of the flash, with every block immediately accessible.

01 / The Problem

A nation buys its most expensive media to hold its coldest data.

A sovereign AI program has two scarce assets, GPUs and storage. Both have to run at high utilization for national AI to be affordable at scale. Wand's backend already attacks the compute half. Storage is the other half of the bill, and it is still bought the old way.

AI storage is provisioned for its peak, so nearly all of it is flash, even though, at any given moment, most of the data sitting on it is inactive. The expensive tier runs at low utilization by construction, and capacity bought per workload and per protocol strands in silos that cannot be lent to the workload that actually needs it.

Provisioned for peak
Capacity bought
100% flash
Data active at any given moment
the working set

The expensive tier runs at low utilization by construction. The rest of the flash is holding data nobody is reading.

Bought per workload, per protocol
Workload ABCD

Each array's unused capacity is stranded. It cannot be lent to the workload next to it, so the nation buys its most expensive media again for the next one.

02 / The Solution

Same outcome. Significantly less flash.

StorONE's Real-Time Tiering writes all data to flash at full performance, then continuously identifies inactive blocks and moves only those to high-capacity media inside the same volume: no policies, no bolt-on tools, no restore step, and every block immediately accessible.

Axis 01 · Flash utilization
The hot working set keeps flash economics on the data that earns them.

Everything is written to flash at full performance. Real-time algorithms continuously separate active from inactive data and relocate only the required blocks to high-capacity media within the same volume. The cold majority stops occupying the most expensive tier while remaining directly addressable, with no external system, copy, or recall in the path.

Axis 02 · Capacity utilization
A ministry's unused terabytes are available to the ministry that needs them.

One platform serves all protocols and all media on standard hardware, with each volume carrying its own performance, protection, and retention parameters. Capacity is drawn from a shared pool rather than committed to a dedicated array per workload, and a nation can build that pool from whatever media it can source.

Elective, not required

The Wand backend runs on a program's existing storage where those economics already work. StorONE is brought in where flash cost, footprint, or power is the binding constraint.

Block File Object Any server Any drive
03 / The Impact

Capacity, megawatts, and capital returned to the AI program.

StorONE reports approximately 90% average flash savings at roughly 10% cost-performance and overhead impact, against the 30–40% that legacy data reduction returns for 80–85% overhead.

~90%
Average flash savings
vs. 30–40% from legacy data reduction

At roughly 10% cost-performance and overhead impact. The same AI workloads served on a fraction of the flash.

9x
Up to 9x more value from flash
Validated by leading storage hardware manufacturers

One platform across block, file, and object, on whatever servers and drives a program can source, with no vendor lock-in and no migrations.

Drill-Down / How a block moves

One volume. Two media. No restore step.

Real-Time Tiering operates inside the active production volume. There is no policy engine to tune, no second system to copy into, and nothing to recall. The block's address never changes.

Step 01
Write

Every write lands on flash at full performance. Cacheless, DirectWrite: no ingest penalty, no staging tier.

Step 02
Detect

Real-time algorithms continuously separate active from inactive data at block level. No schedules, no policies to author.

Step 03
Relocate

Only the inactive blocks move to high-capacity media, inside the same volume, transparently, at roughly 10% overhead impact.

Step 04
Serve

Every block stays immediately accessible at its original address. No restore step in the retrieval path.

One volume · after Real-Time Tiering
single address space · no recall
10%
90%
Active blocks on flash, at full performance
Inactive blocks on high-capacity media, immediately accessible

Illustrative distribution. StorONE reports approximately 90% average flash savings across deployments; the split for a given workload depends on its active working set.

Inside the engine

Four mechanisms, none of them in the data path.

Cacheless DirectWrite ingest

Writes commit straight to flash. No write cache to warm, flush, or lose, so ingest performance is the same on block one and block one billion, and there is no staging tier to size.

Block-level activity map

The engine tracks access at block granularity, continuously: not files, not LUNs, not a nightly scan. An administrator authors no policies and sets no schedules.

Transparent relocation

Only inactive blocks move, and they move within the volume rather than into a second system. Relocation runs as background work at roughly 10% cost-performance and overhead impact.

Single address space

A relocated block keeps its address. Nothing above the volume (no application, agent, or retrieval pipeline) is aware that the media underneath it changed.

The platform stack

Storage services, separated from hardware.

Real-Time Tiering is one layer of a single platform. The Virtual Storage Container™ abstracts the hardware beneath it, so the same services run on whatever servers and drives a nation can actually source.

Protocol layer
Block · FC iSCSI NVMe-oF File · NFS SMB Object · S3
One-Volume
ONE-Volume Architecture
Each volume is an independent storage system, carrying its own protocol, media, protection, and retention parameters.
Data services
vRAID vReplicate Immutable snapshots Encryption MFA · multi-admin auth
Tiering
Real-Time Tiering engine
DirectWrite ingest, block-level activity map, transparent relocation, inside the volume, with no restore step.
Abstraction
Virtual Storage Container™
The hardware-abstraction layer that separates every service above from the media and servers below.
Media
NVMe flash SSD HDD Cloud
Hardware
Any server. Any drive. Any cloud.
One-Volume Architecture

One pool. Per-volume parameters.

Because each volume is an independent storage system, a ministry's training corpus, another's archive, and a third's inference cache can share one pool of capacity while keeping different protocols, protection, and retention, instead of each buying its own array.

volume
protocol
media
protection
retention
training-corpus
S3
Flash + HDD, RTT
vRAID
7 years
inference-cache
NVMe-oF
All flash
vRAID
30 days
ministry-records
SMB
Flash + HDD, RTT
Immutable snapshots
10 years
agency-dr
NFS
HDD + cloud
vReplicate, async
Indefinite

Illustrative volumes. Every row draws from the same shared capacity pool on the same standard hardware; the parameters are per-volume, not per-array.

What does not change when a block moves

The same enterprise storage services are available across all tiers, so tiering is not a trade against protection or recoverability.

Address
Unchanged: no recall, no restore
Snapshots
Immutable, across all tiers
Replication
vReplicate, any-to-any
Encryption
Applied on every tier
Placement
Where StorONE sits in the sovereign stack
Wand AI
Governance & AI labor
Country-level governance, the hybrid human-AI OS, certified AI labor, and autonomous agents.
Wand AI
Sovereign compute base
Every data owner runs on shared capacity rather than a dedicated carveout, raising utilization from 25% to more than 80%.
Ecosystem
Inference privacy, confidential computing, model robustness
Independent capabilities. Each can be adopted on its own; none is a prerequisite for another.
StorONE
Storage efficiency layer
Real-Time Tiering and the ONE-Volume Architecture, on any server and any drive, across block, file, and object.

The capabilities are independent, but they move the same number: cost per outcome for every ministry and agency the sovereign backend serves.

Better Together

Utilization on both sides of the stack.

Wand AI brings

The sovereign infrastructure for AI labor: an operating system through which ministries, public institutions, and critical national organizations deploy AI agents alongside humans, with identity, objectives, authority, security, governance and auditability built in, plus sovereign control over models, compute, data, policies and operations.

Production-scale deployments with leading banks, asset management firms, hedge funds, consulting firms, and system integrators. Founded 2023, headquartered in Palo Alto with offices in Palo Alto, New York and Abu Dhabi.

StorONE brings

Up to 9x more value from flash through a single, hardware-agnostic enterprise storage software platform. Patented Real-Time Tiering manages data across flash and HDD within the same active production volume, keeping all data online and immediately accessible.

The same enterprise services, including snapshots, replication, data protection and encryption, across all tiers, on block, file, and object, with the freedom to choose hardware on price and availability.

The announcement · News · AI Technology

Wand AI adds StorONE as a Storage Efficiency Layer to Enable Sovereign AI on High-Utilization Infrastructure

StorONE's Real-Time Tiering technology becomes an additional storage efficiency capability within Wand's sovereign AI labor infrastructure.

PALO ALTO, Calif., UNITED STATES, August 18, 2026 — Wand AI today announced the addition of StorONE, the leader in Real-Time Tiering architecture, to its Sovereign AI offering, giving sovereign programs a new option for raising the utilization of their storage, as the Wand backend already does for their compute. By incorporating StorONE's Real-Time Tiering into the Wand ecosystem, programs that elect it can serve the same AI workloads on a fraction of the flash they would otherwise have to buy without losing performance or immediate access to any of their data.

A sovereign AI program has two scarce assets, GPUs and storage. Both have to run at high utilization for national AI to be affordable at scale.

National AI infrastructure is judged on cost per outcome, and cost per outcome is a utilization problem. Wand's sovereign backend already attacks the compute half: by letting every data owner run on shared capacity rather than a dedicated carveout, the architecture is designed to raise utilization of national compute from roughly 25% to more than 80%. Storage is the other half of the bill, and it is still bought the old way.

However, AI storage is provisioned for its peak, so nearly all of it is flash, even though, at any given moment, most of the data sitting on it is inactive. The expensive tier runs at low utilization by construction. Capacity compounds the problem. Bought per workload and per protocol, it strands in silos that cannot be lent to the workload that actually needs it. A nation ends up buying its most expensive media to hold its coldest data, and buying it again for the next workload.

The existing storage market does not resolve this. Data reduction returns roughly 30% to 40% capacity savings while consuming 80% to 85% of CPU and memory overhead — paying for storage efficiency with the compute the AI stack was built to use. Archive tiers lower the bill by putting cold data behind a restore step, which is the one thing a retrieval workload cannot tolerate. And proprietary arrays tie capacity to hardware only one vendor can supply, at a moment when sovereign programs need to buy the drives that are actually available to them.

The addition of StorONE closes that gap. The company's Real-Time Tiering writes all data to flash at full performance, then continuously identifies inactive blocks and moves only those to high-capacity media inside the same volume — no policies, no bolt-on tools, no restore step, and every block immediately accessible. StorONE reports approximately 90% average flash savings at roughly 10% cost-performance and overhead impact, against the 30% to 40% that legacy data reduction returns for 80% to 85% overhead. Because the ONE-Volume Architecture separates storage services from hardware, any server, any drive, across block, file, and object, capacity is pooled across workloads instead of stranded in per-workload arrays, and a nation can build that pool from whatever media it can source. StorONE is elective rather than required: the Wand backend runs on a program's existing storage where those economics already work, and StorONE is brought in where flash cost, footprint, or power is the binding constraint.

"Nations are sizing AI programs around GPU utilization and then quietly losing the same money one layer down, on flash that is mostly holding cold data," said Gal Naor, Founder and CEO of StorONE. "Real-Time Tiering moves only the inactive blocks to high-capacity media inside the same volume, so everything stays immediately accessible. One of our enterprise customers went from nine cabinets to two, resulting in an 80% smaller footprint and 78% drop in power usage. At national scale, that is capacity, megawatts, and capital returned to the AI program."

"Sovereign AI is an economics question before it is a technology question, and the answer is utilization on both sides of the stack. We are pleased to welcome StorONE into Wand's sovereign technology ecosystem, adding a storage efficiency capability that ministries, institutions, and agencies can draw on where flash economics are the constraint, serving the same workloads on far less flash, and spending the difference on the AI labor itself," said Cristian Felix, Chief AI Architect of Wand AI.

How It Works

Utilization is raised on two axes. Flash utilization. Everything is written to flash at full performance; real-time algorithms continuously separate active from inactive data and relocate only the required blocks to high-capacity media within the same volume. The hot working set keeps flash economics on the data that earns them, and the cold majority stops occupying the most expensive tier while remaining directly addressable, with no external system, copy, or recall in the path. Capacity utilization. One platform serves all protocols and all media on standard hardware, with each volume carrying its own performance, protection, and retention parameters. Capacity is drawn from a shared pool rather than committed to a dedicated array per workload, so a ministry's unused terabytes are available to the ministry that needs them.

The effect compounds beyond the storage line item. Less flash for the same workload means fewer racks, less power, and less cooling in one enterprise deployment, a reduction from nine cabinets to two, with an 80% decrease in data center footprint and a 78% reduction in power consumption. In a national AI build, those are the constraints that determine how much GPU a program can actually stand up.

The capability is complementary to the inference privacy, confidential computing, and model robustness layers already in Wand's sovereign ecosystem. Those raise the utilization and trustworthiness of national compute; StorONE raises the utilization of the storage underneath it. The capabilities are independent, each can be adopted on its own, and none is a prerequisite for another, but they move the same number: cost per outcome for every ministry and agency the sovereign backend serves.

For more information

To learn more about StorONE and Wand AI, please contact us.

~90% flash savings Up to 9x more value from flash Any server. Any drive.

Request a Demo