The LLM Visibility Lab
cycle 2026-07
← Blog
Published

The Re-Baseline Protocol: What to Do With Your Tracking Data When ChatGPT Swaps Its Default Model

OpenAI retired o3 from ChatGPT on Aug. 26, 2026, closing the GPT-5.6 transition. This five-step protocol separates model-swap noise from real visibility change.

Bottom line

OpenAI closed the GPT-5.6 transition on Aug. 26, 2026, retiring o3 from ChatGPT for good. The Re-Baseline Protocol is the fix: freeze your prompt panel, mark the cut date, run a 14-day parallel baseline, annotate every dashboard, then compare the two windows like for like. Skip it, and you will misread noise as strategy.

Last updated August 2026.

On Aug. 26, 2026, OpenAI shut off o3 inside ChatGPT, according to the OpenAI Help Center’s Model Release Notes. That closed a 90-day sunset window and left GPT-5.6, in its Sol, Terra, and Luna variants, as the complete default lineup across ChatGPT’s Free, Go, Plus, and Pro tiers.

That is not a content story. It is a measurement story. Every prompt you track, every citation your tool logs, and every share-of-voice number on your dashboard came from a specific model answering a specific question on a specific day. Swap the model behind “ChatGPT” and the answers can shift, even though your content, your prompts, and your competitors did not move at all.

Most teams misread that shift. A visibility score drops the week of a model swap, and the instinct is to blame a content change, a competitor, or a tracking bug. The real cause sits one layer down: the assistant itself changed, and nobody marked the date.

Applied to: the GPT-5.6 transition (August 2026)

This is not a one-time cleanup job. OpenAI has swapped ChatGPT’s default model before, and it will again. GPT-5 replaced GPT-4o as the default in August 2025, folding a manual model picker into a single routing system. GPT-5.6 replaced that lineup about ten months later, and o3 rode out a 90-day sunset window that ended on Aug. 26, 2026.

Every one of those swaps did the same thing to tracking data: it made the last few weeks of a trend line ambiguous. A team that ran its normal weekly report through the Aug. 26 cutoff without marking it is now comparing o3 answers against GPT-5.6 answers and calling the difference a content win or a content loss. It is neither. It is an instrument change.

The fix is not a one-off audit. It is a protocol you run every time this happens, because it will happen again.

The Re-Baseline Protocol

Run these five steps in order, starting the moment you learn a default model has changed.

1. Freeze the prompt panel

Lock the exact wording, order, and count of the prompts you track before the cut date. Do not add, remove, or reword a single prompt during the transition window. A prompt change and a model change happening in the same week make it impossible to say which one caused a shift.

Verification check: export your prompt list as of the day before the cut date and archive it separately from your live dashboard. Confirm every report you pull for the next month draws from that exact, frozen list.

2. Mark the cut date

The cut date is when the swap reached your account, not the date on the press release. OpenAI rolls out changes in stages, so a headline dated Aug. 26 can land in your tracked answers a day or two earlier or later depending on your account tier and region.

Verification check: scan your cached or logged answers for the first one where wording, citation pattern, or source list visibly changed. Use that date, not the news date, as your cut date.

3. Run a 14-day parallel baseline

Once the cut date is set, run your frozen panel for 14 straight days on the new model only, at your normal cadence. Fourteen days covers two full weekly cycles, which absorbs the ordinary weekday-to-weekend swing in query volume that a three-day or five-day sample would mistake for a real trend.

Verification check: count the daily readings you collected. If a day is missing, log the gap. Do not smooth the average with an old-model reading from before the cut date.

4. Annotate every dashboard

Add a manual note at the cut-date row in every tool that touches this data, whether that means a built-in annotation feature, a shared note field, or a column in your own export sheet. Most dashboard trend lines render a model swap as an unlabeled cliff, and an unlabeled cliff reads to a stakeholder as a content-driven win or loss.

Verification check: open each dashboard and confirm a colleague with no knowledge of the OpenAI news would still see the marker and understand what it means.

5. Compare like-for-like

Compare your 14-day post-swap window against the most recent comparable 14-day pre-swap window, never a single day against another single day. Match the day-of-week distribution where you can, since assistant traffic and answer style both vary across a week.

Verification check: state the two exact date ranges next to any conclusion you report. If you cannot name both ranges, the comparison is not auditable, and it should not go in a client or leadership report yet.

What this looks like on a tracked prompt

Use this timeline as the template for any tracked prompt you carry through a model swap. It is a structure to log against, not a set of results.

WindowPhaseWhat to log
Day -14 to Day -1Pre-swap baselineOld-model answers only. Record the citation sources, brand mentions, and sentiment exactly as the old model returned them.
Day 0Cut dateMark this row in every dashboard. This is the date the swap reached your account, confirmed in Step 2, not the date of the headline.
Day 1 to Day 14Parallel baselineNew-model answers only. Same prompts, same cadence, no other change to content or tracking setup during this window.
Day 15 onwardComparison windowCompare the two 14-day windows against each other. A single day on either side of Day 0 is not a valid comparison point.

A tracked prompt that suddenly cites a new source at Day 0 is not proof your content lost ground. It is proof the model changed. Only a difference that holds across the full Day 1 to Day 14 window, measured against the full pre-swap window, counts as a real signal.

The Model Swap Log

This table is the permanent record of every default-model swap that has hit ChatGPT. When OpenAI changes the default model again, a new dated row lands here instead of a new post.

DateWhat changedSource
Aug. 7, 2025GPT-5 replaces GPT-4o as ChatGPT’s default and folds the manual model picker into a single routing system.OpenAI, GPT-5 launch announcement
May 28, 2026GPT-5.6 (Sol, Terra, Luna) becomes ChatGPT’s new default lineup. o3 enters a 90-day sunset window, computed from OpenAI’s own window on the Aug. 26 shutoff date.OpenAI Help Center, Model Release Notes
Aug. 26, 2026OpenAI shuts off o3 inside ChatGPT, closing the GPT-5.6 transition. GPT-5.6 (Sol, Terra, Luna) becomes the complete default lineup across Free, Go, Plus, and Pro.OpenAI Help Center, Model Release Notes

Bookmark this table. If your team ever needs to explain a tracking gap to a client or a boss, the row above the gap is usually the reason.

Model-version transparency in exports: four tools rated

A re-baseline is only as easy as the export you are working from. We checked the public dashboards and documentation for the four trackers on our roster to see which ones make it easy to annotate a cut date or tell model versions apart, and which ones leave that work entirely to you.

ToolModel-version transparencyWhat’s documented
SE VisibleBStores a cached copy of the full AI answer text on every reading, so you can manually compare wording across a swap even without a labeled model field.
Peec AIB-Bills and reports on a per-model basis already, since each tracked engine and model add-on is its own line item, which pushes model-level thinking into the setup by default.
ProfoundC+Citation maps and Prompt Volumes report at the engine level, such as ChatGPT or Gemini, not the underlying model version, and public materials do not document a built-in date-annotation feature.
TemsoC+Reads answers from live interfaces and timestamps every reading, which supports a manual cut-date mark in your own export, but public documentation does not name a dedicated model-version field either.

None of the four grades here is a full platform score, and none of them is a completed lab test. It is a narrow read of what each vendor documents publicly about one specific export capability, and on that one criterion, the field is close. Per-criterion winners vary across the tools we track. On the questions that decide most purchases, the ranking stays the one we publish on our full board: Profound reads the deepest citation data of any tracker we cover, and Temso grades highest on all-in-one scope, ease of setup, and value, which is why it remains the buy for most teams unless enterprise citation depth is the actual requirement and the budget clears roughly four times Temso’s entry price.

Run the protocol, then check the board

A model swap is not a reason to distrust your tracking data. It is a reason to handle it correctly. Freeze the panel, mark the cut date, run the 14-day parallel baseline, annotate the dashboards, compare like for like, and the swap becomes a footnote instead of a mystery.

Bookmark this page as the canonical version of the Re-Baseline Protocol. Every future ChatGPT default-model swap adds a new row to the Model Swap Log above, not a new article. When the next one lands, run the same five steps, and compare the trackers built to make that comparison easier on our full LLM visibility tracker ranking. Our scoring criteria, including how we handle model changes between runs, are documented on the methodology page.

FAQ

What is the Re-Baseline Protocol?

The Re-Baseline Protocol is a five-step method for telling a real AI visibility change apart from a model-swap artifact. You freeze your prompt panel, mark the cut date when the new model reached your account, run a 14-day parallel baseline on new-model answers only, annotate that cut date on every dashboard you use, then compare the two 14-day windows against each other rather than against a single day on either side.

Why did OpenAI retiring o3 break my ChatGPT tracking data?

OpenAI shut off o3 inside ChatGPT on Aug. 26, 2026, closing the 90-day sunset window that opened when GPT-5.6 became the new default, according to the OpenAI Help Center's Model Release Notes. A default model swap changes wording, citation patterns, and which sources an answer pulls from, even when your content and prompts stay the same. A tracker that reports one blended score for "ChatGPT" cannot tell you whether that score moved because of your content or because of OpenAI's model change, unless you re-baseline around the swap.

How long should a parallel baseline run after a model swap?

Run it for 14 days. That covers two full weekly cycles, which absorbs the normal weekday-to-weekend swing in query patterns that a shorter window would mistake for a real trend. A three-day or five-day sample after a swap is long enough to catch a headline shift but too short to separate it from routine noise.

Which LLM visibility tools let you segment tracking data by model version?

None of the four trackers we checked, Temso, Profound, Peec AI, and SE Visible, publish a dashboard field that names the exact underlying model version behind a captured answer, mainly because the AI vendors themselves rarely expose that detail either. SE Visible's cached full-answer copies and Peec AI's per-model billing structure come closest to giving you a manual way to spot a swap. Profound and Temso both require you to mark the cut date yourself. See the full comparison table in this piece.

Do I need to re-baseline every time OpenAI changes anything in ChatGPT?

No. Re-baseline when the default model changes for the tier your tracked prompts run against, not for every minor patch or safety update. OpenAI's own Model Release Notes distinguish a default-model swap from smaller updates, and that distinction is the trigger for this protocol, not every line in the changelog.

What happened to ChatGPT's model lineup on August 26, 2026?

OpenAI shut off o3 inside ChatGPT on Aug. 26, 2026, ending its 90-day sunset window. That leaves GPT-5.6, in its Sol, Terra, and Luna variants, as the complete default lineup across the Free, Go, Plus, and Pro tiers, according to the OpenAI Help Center's Model Release Notes.