Data Orchestrator

Move data to the compute

Limestone Data Orchestrator helps AI labs move datasets, model weights, containers, artifacts, and logs across fragmented storage and GPU compute environments with high throughput, reliable execution, and clear monitoring.

Data Orchestrator jobs

Control plane view

healthy
hydrate-model-weightsRUNNING
82% complete20.1 GB/s
return-training-artifactsRUNNING
47% complete14.7 GB/s
cleanup-ephemeral-cachePENDING
0% complete-

12.8 PB

found

9.4 PB

copied

2

errors

Problem

Data and compute no longer live in the same place.

AI workloads span clouds, regions, and private GPU clusters, while the data they need is scattered across cluster storage, object stores, file systems, databases, lakes, warehouses, and registries.

Teams connect these systems with fragile scripts and cloud-specific pipelines. Slow or silent failures leave accelerators idle and pull platform engineers away from infrastructure work.

Solution

Data movement jobs for AI workloads.

Limestone Data Orchestrator is a vendor-neutral layer for defining, tracking, and running data movement jobs across varied storage environments. Its control plane manages state, progress, and errors while the data plane executes transfers near compute environments.

It delivers fast, scalable, observable, reliable, and secure petabyte-scale movement across private clusters and public clouds. This keeps customers in control of their data while reducing accelerator idle time.

Product

Data Orchestrator is built for speed, reliability, and monitoring from the first job.

Data Orchestrator makes data movement fast, reliable, observable, and easier to operate as AI infrastructure becomes more distributed.

It is designed for the workflows between storage and compute: hydrating pre-training clusters, keeping post-training loops synchronized, placing complete model bundles near inference capacity, and cleaning up ephemeral environments.

High

throughput data transfer

Data Orchestrator achieves 95%+ network saturation and up to 10× higher throughput than existing tooling by scaling the data transfer horizontally.

Technical detail+

Performance scales horizontally by decoupling the control and data planes, using stateless workers deployed close to storage and compute clusters. Distributed, pipelined, and parallelized fan-out execution eliminates noisy neighbor effects and coordination bottlenecks. Streaming I/O with backpressure handles both small and large files efficiently, and zero-copy checksums provide end-to-end data integrity.

Durable

job execution

Accepted jobs are persisted, decomposed into retry-safe work, and tracked through monotonic state transitions.

Visible

operations

Progress counters, job state, manifests, and structured errors make long-running movement easy to monitor and reason about.

A job model for infrastructure teams

Copy workflows are submitted as durable jobs with explicit state, counters, manifests, and structured errors. The CLI and Python SDK interface with one unified control plane across storage and compute resources for humans and their agents.

do job copy \
  --origin s3://training-data/frontier-v4 \
  --destination cluster://h100-pool-a/datasets \
  --follow-symlinks=false

job_id: job_01HX9B7M6R
state: RUNNING
throughput: 20.1 GB/s
files_copied: 18,402,117

Use cases

Keep GPUs busy with full workload portability.

01

Pre-training

Hydrate base-model artifacts, containers, configuration, and datasets on the cluster selected for a run. Make the complete dependency set available to the GPUs, then return checkpoints, logs, and evaluation outputs to durable storage.

02

Post-training

Keep iterative fine-tuning and reinforcement-learning loops supplied with task environments, trajectories, rewards, checkpoints, and evaluation results. Synchronize updated weights to rollout fleets and move work between training and generation pools without losing recovery state.

03

Inference

Stage complete serving bundles during autoscaling, failover, and cold starts; place new model versions with their dependencies while keeping prior versions ready for rollback; and move batch inputs to available compute pools before returning outputs to durable storage.

04

Agentic workflows

Use local copies of data in isolated sandboxes so each agent can work quickly and safely. Then use the CLI to synchronize changes back to central storage to maintain a shared, durable source of truth.

Results

More utilization. Faster research. Faster releases.

Infrastructure efficiency

Increase GPU-fleet utilization and reduce wasted GPU-hours and spend.

Researcher velocity

Shorten time from job submission to execution start, returning valuable time to researchers.

Engineering capacity

Free platform engineers from custom pipeline work so they can focus on infrastructure.

Business impact

Speed model iteration and bring model releases forward.

FAQ

Designed for infrastructure reality.

What problem does Data Orchestrator solve?+

AI workloads often run far from the data they need, leaving expensive GPUs idle while engineers move data by hand. Data Orchestrator enables your team to move large datasets, models, and other artifacts durably to available compute, then deliver outputs to your desired long-term storage, unblocking your training and inference workloads. Our solution gives AI researchers and engineers time back while freeing up precious GPU cycles across the fleet.

Who is Data Orchestrator for?+

Data Orchestrator is built for AI teams working with diverse compute and storage systems across clusters, regions, and clouds. It helps researchers move faster with their experiments while giving infrastructure teams a reliable, consistent way to manage data across training, inference, and agentic workloads.

How does Data Orchestrator help researchers and engineers?+

Researchers get faster access to the datasets, checkpoints, model weights, containers, and artifacts their workloads require. Infrastructure engineers get durable data ingestion, fine-grained observability, and a consistent way to stage inputs and collect outputs. The result is faster experimentation and less compute time lost waiting for data.

How does Data Orchestrator enable workload portability?+

Data Orchestrator separates data movement from the underlying compute and storage environments. Workloads can easily move across GPU clusters while retaining a consistent view into the data, allowing teams to use the compute that best meets their performance, availability, residency, or cost requirements.

Does Data Orchestrator access or store my data?+

Data Orchestrator moves data only between the storage and compute environments you authorize. In the managed offering, DO encrypts credentials, stores them securely, and uses them only for authorized operations. BYOC means Bring Your Own Cloud: data movement execution remains within your environment, preserving your data-residency and security boundaries. We integrate with Kubernetes and will support Slurm, SkyPilot, or SUNK runtimes in the future.

What storage systems are supported?+

We are committed to making the Data Orchestrator storage agnostic, so your team can use whichever storage systems best complement your workloads. Today, we support object storage (Amazon S3, Google Cloud Storage, and Azure Blob Storage) and file systems (NFS, Lustre, VAST, and CephFS). We're working on support for data lakes, data warehouses, container registries, and model registries.