Low-code fine-tuning for LLMs

Small models,frontier results.

The end-to-end platform for small language models. Tuned to your task, a small open model matches frontier accuracy at a fraction of the cost. Own your intelligence: private, compatible, no lock-in. Ever.

No credit card · pay only for GPU minutes you use$0.03 / GPU·min flat
  • Bring data your way: CSV import, API, or the built-in chat editor
  • OpenAI-compatible API: swap one base URL and keep your code
  • Weights are yours to download anytime. No lock-in.
Small models beat giants on your task~1/10 the costFaster responsesYour data trains only your models

From raw data to serving API

One managed pipeline. You bring examples; we handle GPUs, training and serving.

Your dataCSV · API · UI
  1. DatasetsUpload CSV, push via API, or write examples in the editor
  2. Fine-tuneOne-click training on small open models
  3. EvaluateLive loss curves and checkpoints while it trains
  4. ServeHosted inference in our cloud, or download the weights

Managed GPU infrastructure · $0.03 per GPU minute · stop anytime

PredictionsREST API · OpenAI-compatible

// do the math

Estimate your savings

Compare what you pay a big-model API today against a small model tuned to the same task, served on our GPUs.

Use case

Commercial API today

Estimated cost

Unavailable

Assumes one always-on GPU at the rate above serving a tuned small open model at the selected use case. Training itself is separate and one-off. Scale-to-zero, bursty traffic or a bigger model change the math; this is the list-price estimate. 8/3/2026.

// small beats big

A small model tuned to your task beats a giant generalist

Case-based research on fine-tuned open models keeps finding the same result: you do not need a bigger model. You need a model that knows your data.

Accuracy

Frontier quality, on your task

Fine-tuned on your examples, a small open model matches or beats models many times its size, on the one task that matters to you.

Cost

~1/10 the cost per answer

Fewer parameters need a fraction of the GPU. One rented replica serves a steady stream of requests for a flat per-minute rate, instead of a per-token bill that grows with every answer.

Latency

Milliseconds, not seconds

A tuned small model decodes at about 25 ms per token on our production hardware, so a short answer is back before a big model has finished thinking about it.

Why teams pick ReOpenly

Keep the code you already wrote

Point your OpenAI SDK at a new base URL and you’re done. Works with LangChain, n8n and friends.

const baseURL = "https://reopenly.com/api/projects/me/proj/openai/v1";
const client = new OpenAI({ baseURL, apiKey: "roy_..." });

Your traffic becomes training data

Opt in to auto-save a share of inference requests to your dataset, curate them, and fine-tune. Your data only ever trains models you own.

Leave whenever you like

Every fine-tuned model can be downloaded anytime. Serve with us because it’s easy, not because you must.

Pricing you can read in one line

$0.03 per GPU minute, training or inference. 1 credit = $1. That is the whole pricing page.

Pay as you go

$0.03 / GPU minute

Training or inference, same flat rate. Paid inference has a one-second minimum.

Welcome bonus

10 free credits

Every new account starts with 10 credits, enough for several fine-tunes on small models.

1 credit = $1. No tiers, no seats, no hidden fees.

Estimate your cost →

Your first fine-tune is on us.

From CSV to a deployed, OpenAI-compatible endpoint in under an hour.

No credit card · weights yours to download