Refactr

Put private AI to work.

Refactr deploys private AI assistants inside your infrastructure. Use the hardware you own, or let us source the system you need.

Your data is too valuable to send away.

Your team wants to use AI. Your company cannot send sensitive prompts, documents, or outputs to public AI services. That leaves employees working around the policy, or without AI at all.

Documents leaving a secure container toward a public cloud service.

Bring AI to your data.

Refactr deploys the right model on infrastructure you control. Your team gets fast AI for everyday work, while your prompts, documents, and outputs stay inside your environment.

Documents, AI chat, and local hardware connected in a closed loop inside one environment.

What we deploy

Private AI for work your team already does.

Start with a document or coding assistant, deployed inside your own environment. More workflows follow the same path.

Document assistant

Ask questions, summarize material, draft from documents, and work through sensitive information without sending it to a public AI service.

Summarize the open decisions

Your company stays in control.

You decide where the system runs, which model it uses, and what data it can access. Refactr deploys and maintains the stack.

You set the boundaries

Choose the environment, approved data,
users, and model.

Work stays inside

In a private deployment, company work stays inside the environment you choose.

Choose or fine-tune the model

Deploy a compatible open model as-is, or fine-tune one for your workflow and approved data.

Use or source the hardware

Start with the GPUs you have. If more is needed, we can specify, procure, and supply it.

Refactr runs the stack

We handle setup, monitoring, updates, performance tuning, and ongoing support.

How Refactr runs the stack

Get more from local GPUs.

We built Epsilon, our inference layer, to route requests across local engines. It improves response time and hardware use without sending work outside your environment.

A compact inference engine connecting two local compute rails into one efficient output path.

With Epsilon vs Basic load-based routing

TTFT p95

Tail time to first token on the same dual warm pool.

With Epsilon

0 ms

Without Epsilon

0 ms

0.0% lower tail TTFT

GPU-s / valid

Active GPU-seconds per valid request on the same pool and trace.

With Epsilon

0.000

Without Epsilon

0.000

0.0% lower GPU-s per valid request

End-to-end p50

0.0% lower E2E p50. With Epsilon vs ordinary routing. Shorter bar is faster.

Under load

E2E p50 as concurrency increases. 100% completion both sides.
Ordinary routing
With Epsilon

~0% cheaper per million output tokens

With Epsilon costs about 0% less per million output tokens than Without Epsilon at every GPU rate we modeled. Throughput on the same pool is 0.0% higher. Dollar rates below are approximate scenarios, not invoices.

Without Epsilon
With Epsilon

Where Epsilon fits.

Epsilon is built for efficient local inference on infrastructure you control. Other tools serve different operating models.

Swipe to compare systems

Axis
Epsilon
LiteLLM
llm-d
Dynamo
Routes across local GPU engines
YesWarm llama.cpp + vLLM
NoProvider / API backends
PartialvLLM-centered
PartialNVIDIA engines only
No cluster required
YesBare-metal / small host
YesProxy process
NoCluster control plane
NoDC orchestrator
Benchmarks configs before routing
PartialMeasured curves + policy
NoNo recipe gate
NoNo recipe gate
NoNo recipe gate
Single consumer GPU / small fleet
YesBuilt for one box to small fleet
PartialWorks; not fleet-native
NoCluster-scale fleets
NoNVIDIA datacenter

A recipe is the engine, model, and quantization each worker runs.

Product fit from public docs. Latency on this hardware was measured only for Epsilon vs basic load-based routing above.

Same two-GPU pool (2× RX 7900 XTX).

How we start

Start with one real job.

Don’t roll AI out across the company on day one. Pick one recurring task, put it in front of the people who do it, and see if it earns a wider rollout.

  1. Find the job

    Choose work people repeat often enough for an improvement to matter.

  2. Draw the line

    Decide what it can see, what it cannot, and who can use it.

  3. Put it in people's hands

    We set it up in your environment and support a small group as they start using it.

  4. Decide what to do next

    If it is useful, expand it. If it needs work, improve it. If it is not worth it, stop.

FAQ