
In Short: Which Platforms Offer Reinforcement Fine-Tuning as a Managed Service?
As of September 2026, three major clouds document reinforcement fine-tuning (RFT) as a managed service: Microsoft Foundry (o4-mini, plus invitation-only GPT-5), Amazon Bedrock (Nova 2 Lite, gpt-oss-20b and Qwen3 32B), and Google's Gemini Enterprise Agent Platform, formerly Vertex AI (Gemini 3.5 Flash and Gemini 3.1 Flash-Lite, in preview). OpenAI's own API is closed to new fine-tuning customers, and Microsoft's Frontier Tuning is in private preview.
All three work the same way at heart. You supply prompts and a way to score answers, and the service trains the model towards higher scores. The differences sit in which models you can tune, how you write the grader, how you pay, and what you are allowed to put in the training data.
RFT vs Supervised Fine-Tuning
Supervised fine-tuning (SFT) teaches by example. You provide input-output pairs - the question and the ideal answer - and the model learns to imitate them. It works well for format, tone and narrow tasks, but it needs a lot of well-written answers, and the model only learns what your examples show.
Reinforcement fine-tuning teaches by feedback. You provide prompts plus whatever reference material a grader needs. The model produces several candidate responses, a grader scores each one, and training nudges the model towards the higher-scoring behaviour. Microsoft describes it as training "through a reward-based process, rather than relying only on labeled data".
That makes RFT a good fit when judging an answer is easier than writing the perfect one: extracting the right clause, reaching the correct classification through several reasoning steps, producing output that passes a schema and a set of business rules. It is a poor fit when "good" cannot be stated precisely, because a vague grader gets exploited. The Foundry documentation warns specifically about reward hacking, where training and validation rewards drift apart because the model has found a way to score well without doing the task.
If you want the organisation-wide version of this idea, Microsoft's Frontier Tuning approach applies reinforcement learning to workflows inside your compliance boundary. We cover it separately. This post is about the self-service RFT jobs you can run today.
Microsoft Foundry (Azure OpenAI)
- Models: o4-mini (2025-04-16), generally available. GPT-5 (2025-08-07) is generally available but gated, by invitation only through your Microsoft account team.
- Graders: string check (exact or contains match), text similarity (fuzzy match, BLEU, ROUGE and others), model graders that score with a prompt, Python code graders running in a sandbox with no network access, and a multigrader that combines several scores with an arithmetic formula. Endpoint graders that call your own API are in private preview.
- Data: JSONL in chat completions format, with the final message from the user. Extra fields carry the ground truth your grader reads. Separate training and validation files are required.
- Pricing: billed on core training time, not tokens. Microsoft's cost guide uses $100 per training hour for o4-mini as its worked example, with model-grader tokens billed separately. Jobs pause automatically at $5,000 of training and grading spend. Deployed models carry an hourly hosting fee on Standard deployments.
Foundry is the natural home if your agents already run there - see our guide to choosing models in Microsoft Foundry. One thing to know: OpenAI is winding down its own fine-tuning platform. Organisations that have never fine-tuned there can no longer create jobs, and from 6 January 2027 existing customers cannot create new ones either. Microsoft still documents o4-mini RFT as generally available in Foundry, so if you need OpenAI reasoning models with RFT, Foundry is the route to check.
Amazon Bedrock
- Models: Amazon Nova 2 Lite (us-east-1), and the open-weight gpt-oss-20B and Qwen3 32B (us-west-2). Bedrock also offers OpenAI-compatible fine-tuning APIs for the open-weight models.
- Graders: AWS Lambda functions for rule-based scoring, or a model-as-a-judge set up in the console, which Bedrock turns into a Lambda for you. Training uses Group Relative Policy Optimization (GRPO).
- Data: JSONL in OpenAI chat completion format, with a reference answer field, and a minimum of 100 records. You can also train directly from existing Bedrock invocation logs, filtered by request metadata. Nova jobs accept up to 20,000 prompts.
- Pricing: hourly training. Bedrock lists $80 per training hour for gpt-oss-20b and Qwen3 32B. The tuned model then runs on on-demand inference at base-model token rates, or on Provisioned Throughput.
AWS's best-practice advice is useful on any platform: test the base model first. If rewards are near zero, do supervised fine-tuning first. If they are above 95 percent, you probably do not need RFT.
Google Gemini Enterprise Agent Platform (Vertex AI)
- Models: Gemini 3.5 Flash and Gemini 3.1 Flash-Lite. Tuning runs in us-central1 or europe-west4, and tuned models are served from the US or EU multi-region endpoint.
- Status: Preview. The documentation labels it a Pre-GA offering and says not to use proprietary, sensitive or confidential data, and not to use it for production.
- Rewards: string matching, a Gemini-based autorater (LLM-as-a-judge), Python code execution, and fully custom rewards hosted on Cloud Run. Up to 16 can be combined with weights. Scores are clipped to the range -1 to 1.
- Data: JSONL in Cloud Storage with a references field for grading metadata. The limit is 5,000 training examples, and text, image, audio and video are supported (video on 3.5 Flash only). You can run RFT on top of an SFT-tuned model.
- Pricing: you pay for tokens generated in the training phase, at Google's published model-tuning rates. From Gemini 3 onwards, a tuned model costs 1.5 times the base model to call.
When RFT Is Worth It
Most teams that ask us about fine-tuning do not need it yet. Here is the order we work through:
- Prompt engineering first. Clearer instructions, examples and structured output fix most quality gaps, cost nothing to train and can be undone in minutes.
- Retrieval next. If the model is missing facts - policies, product data, prices - the fix is retrieval, not training. A well-built Foundry Agent Service agent with good grounding usually beats a tuned model with stale knowledge.
- Supervised fine-tuning when you have hundreds of good examples and need a consistent format or tone, or a smaller, cheaper model to match a larger one on a narrow task.
- RFT when the task needs several steps of reasoning, success can be scored reliably, and prompting has stopped improving. You should also have an evaluation set you trust. Without one, you cannot tell whether a tuned model is actually better.
Do the maths as well. A tuned model means training runs, possible hosting fees and a higher per-call price on some platforms. Weigh that against what you save by moving to a smaller model, using the same inputs as our breakdown of what AI agents really cost to run.
Data Governance Before You Train
RFT has fewer data demands than SFT, but it still bakes your data into a model's weights. Three rules:
- Grader quality is data quality. A grader built on inconsistent reference answers teaches the model inconsistent behaviour, faster. The same discipline applies as in our guide to data quality for AI.
- Check where training runs. Foundry's global and developer training tiers are cheaper but do not guarantee data residency. Google's preview terms rule out confidential data entirely. Bedrock's options depend on region.
- Classify before you export. Prompt logs and invocation logs often contain personal data. Know what sensitive data sits in them before they become a training set - Purview DSPM for AI is where Microsoft estates usually start.
Where Solv Systems Comes In
We help organisations decide whether fine-tuning is the right lever at all, then build the evaluation sets, graders and governance that make the answer measurable. If you are weighing RFT against a better prompt or a retrieval fix, our AI automation consulting team can run that comparison on your own data.
Sources and Further Reading
- Reinforcement fine-tuning - Microsoft Foundry
- Fine-tuning cost management - Microsoft Foundry
- Customize a model with reinforcement fine-tuning in Amazon Bedrock
- About reinforcement learning fine-tuning for Gemini models
- OpenAI API deprecations
- Frontier Tuning: Teaching AI to work the way you do
Frequently asked
Reinforcement fine-tuning (RFT) trains a model against a grader or reward function rather than against labelled answers. The model produces candidate responses, the grader scores them, and training shifts the model towards the behaviour that scores well. It suits tasks where you can reliably judge a good answer even if writing the perfect one is hard.
As of September 2026: Microsoft Foundry (o4-mini generally available, GPT-5 generally available but invitation only), Amazon Bedrock (Nova 2 Lite, gpt-oss-20b and Qwen3 32B), and Google's Gemini Enterprise Agent Platform, formerly Vertex AI (Gemini 3.5 Flash and Gemini 3.1 Flash-Lite, in preview). OpenAI's own API is closed to new fine-tuning customers, and Microsoft's Frontier Tuning is in private preview.
Foundry and Bedrock bill by training time: Microsoft's worked example uses $100 per core training hour for o4-mini, plus grader model tokens, and Bedrock lists $80 per training hour for gpt-oss-20b and Qwen3 32B. Google charges for tokens generated in the training phase, and Gemini 3 tuned-model inference costs 1.5 times the base model price. Always check the current pricing pages.
Less than supervised fine-tuning usually needs, because you supply prompts and grading criteria rather than full ideal answers. Bedrock requires at least 100 training records and recommends starting with 100-200. Google caps RL training datasets at 5,000 examples. Foundry requires separate training and validation files in chat completions JSONL format.
It solves a different problem. Prompting and retrieval fix missing instructions and missing facts, cheaply and reversibly. RFT changes how the model reasons and decides on a narrow, gradable task. Exhaust prompting and retrieval first, and build an evaluation set before considering any fine-tuning.
It depends on the platform and its launch stage. Google's reinforcement tuning is a Pre-GA offering whose terms say not to use proprietary, sensitive or confidential data. On Foundry, global and developer training tiers do not guarantee data residency. Check the terms, training location and hosting option before any regulated data goes near a training job.


