Launch is not the finish line. It’s the handoff.
The industry has spent two decades building a vocabulary that says otherwise. Go-live. Ship date. Launch party. Every one of those phrases frames a relaunch as a thing that ends, and enterprise organizations have organized themselves around that framing so completely that you can read the belief straight off the org chart.
The project team dissolves. The standing meeting comes off the calendar. The budget line closes.
What replaces it is usually nothing.
So quality erodes, and it erodes quietly. Your site doesn’t fall over the week after launch. It drifts: a redirect chain that nobody noticed, a template change that dropped a heading level, an image library that stopped carrying alt text when the new upload flow made the field optional. None of that generates an incident ticket. It accumulates until somebody outside the organization notices first, and by then you’re not managing quality anymore, you’re managing a complaint.
Those aren’t random failures. They’re the accessibility, discoverability, content, and measurement categories in Siteimprove’s Enterprise Redesign and Migration Risk Framework, surfacing in the weeks after launch rather than during the migration everyone planned for.
I’ve watched teams run flawless migrations and lose the site anyway, which is why I think the launch-as-milestone framing does more damage than any technical decision in the project.
This piece argues that the organizations holding quality after a relaunch aren’t the ones with better launches. They’re the ones that converted the project team into a standing function with a named owner and a monitoring cadence before the celebration ended. It moves from what to validate in the first weeks, through how decay shows up and what a standing governance function actually consists of, to the measurement model that tells you whether any of it is working.
The first weeks after launch decide what quality survives
The window right after go-live is when regressions are cheapest to catch and most likely to be missed. Both halves of that sentence matter, and the second one is organizational rather than technical.
Post-launch validation, sometimes called website maintenance or ongoing site monitoring, is the practice of systematically re-checking quality, accessibility, discoverability, and measurement signals against a known baseline in the weeks after a site goes live. It’s not the same as launch QA. Launch QA asks whether the migration worked. Post-launch validation asks whether it’s still working now that real publishing has resumed, real traffic is hitting the site, and real editors are touching templates you tested under laboratory conditions.
Here’s what belongs in that window, and here’s the part most checklists skip: every item needs a name next to it.
- Redirects and link integrity. Broken internal links, redirect chains, and orphaned pages. Whoever owns your URL map owns this one.
- Indexation signals. Crawl coverage, canonical behavior, sitemap submission, robots directives. Google’s site move guidance is explicit that a move completes on a per-URL basis over months, not on launch day, which means your indexation checks have to keep running long after you’ve stopped thinking about the migration. Your technical SEO owner takes it.
- Accessibility regression. Conformance against your stated WCAG target, re-run rather than assumed. The accessibility and discoverability section takes up why these two degrade together. Your accessibility lead owns it.
- Performance baselines. Field data, not lab data, captured as a baseline you’ll measure drift against, covering load time, responsiveness, and uptime. Performance management after launch gets its own section below. Whoever owns your platform performance takes it.
- Measurement continuity. Tag firing, event definitions, goal configuration, data flowing into the reports people actually use. Broken tracking produces improving metrics, which the measurement section takes up in full. Your analytics lead owns it.
An unassigned task in that list is a task that doesn’t happen. You know this. Every enterprise team knows this, but I’ve yet to see a post-launch checklist that shipped with owners already written next to the items, because at the moment it’s written everyone is on the project and ownership feels obvious.
It stops being obvious in about three weeks.
The other correction worth making is about shape. A checklist is an event. What you need is a cadence: the same checks, re-run on a schedule, with thresholds that trigger a response. Siteimprove’s Post-Launch Validation Framework sets out that rhythm across the first 30, 60, and 90 days after migration, and it’s worth working through if you’re standing up a validation schedule from scratch.
Decay announces itself long before anyone reports a problem
Quality erosion is legible. That’s what we get wrong about it. Teams talk about post-launch decay as if it were weather, something that happens to a site rather than something a site displays, and that framing quietly excuses the failure to look.
Decay produces signals, and each one points at a different failure in the publishing process rather than at a single broken thing. The ones that surface first:
- Accessibility errors reappearing in templates that passed at launch, usually where a CMS field went optional or an editor pasted styled markup from another system.
- Impressions and rankings sliding on pages that didn’t change, which is the signature of a structural problem rather than a content one.
- Load-time drift as third-party scripts, unoptimized images, and new embeds accumulate week over week.
- User journeys breaking at the edges: a form that fails on mobile, a filter that returns nothing, a PDF that never got a text layer.
Two definitions worth stating plainly, because they get used loosely.
Accessibility regression is the reappearance of conformance failures on pages or templates that previously met your stated standard. It is a change in state, not a change in standard, and it’s why point-in-time audits mislead: a page that conformed in March tells you nothing about that page in September.
Discoverability regression is the degradation of the structural and technical signals that let search engines and answer engines find, crawl, parse, and represent your content. Broken canonicals, lost headings, removed structured data, and redirect chains all qualify.
Neither of those requires sophisticated detection. Both require somebody to be looking.
So why do they go unseen for months? Not because the signals are subtle, but because detection without an owner is a capability nobody exercises.
The cost of not looking compounds with time, and the compounding is the whole argument. A heading-level regression caught in week two is a template fix. The same regression caught in month nine has propagated across every page built from that template since, and remediating it means touching thousands of pages under time pressure, usually because a complaint or a legal letter set the clock. The work didn’t get harder. The volume did.
If you want a structured way to think about which risks you’re carrying into the post-launch period, the seven-category risk framework is worth reading in full, because the post-launch window is where the accessibility and compliance, performance and discoverability, content, and measurement categories all come due at once. Decide which of those you’re prepared to leave unmonitored before the project team disbands, not after.
Federal accessibility reporting offers a useful external mirror for what happens when programs exist on paper but not in practice. The U.S. General Services Administration runs an annual governmentwide Section 508 assessment, and the FY2025 edition reset its criteria to establish a new baseline, which means you can’t read it as a year-over-year trend. Read it as a snapshot instead. Across 212 agencies and components, conformance averaged 1.96 on a 5-point scale, fewer than half of the most-viewed assets were fully conformant, and roughly half of agencies said they don’t routinely test for conformance as standard practice. Verify these figures against the primary report before publication, since the assessment is reissued annually.
Three findings from that assessment land directly on the argument here. Agencies with dedicated program leadership and clearer management structures showed stronger accessibility integration and better conformance outcomes. Implementation effectiveness rather than agency size was what tracked with conformance. And many agencies were found to hold standalone accessibility policies that never got integrated into operational policy, which limited consistency and enforcement.
Leadership, implementation, and integration all point the same direction: a policy nobody operationalized produces roughly the same outcome as no policy.
If federal agencies can’t close that gap under a statutory mandate, we shouldn’t expect a relaunch to close it on goodwill.
A standing owner is what separates a program from a project
A governance program is not a document. It’s not a committee, a policy PDF, or a slide that says “governance” over a diagram of your CMS.
It’s a person with a calendar and authority.
That’s the uncomfortable version, and I’d rather be blunt about it, because “we need governance” is one of those phrases that can absorb a year of enterprise effort without anything changing. A functioning post-launch governance program has four components, and you can check whether you have them this afternoon:
- A named owner. One person accountable for the state of quality across the estate, with the seniority to make competing teams do work.
- A cadence. Fixed intervals at which the validation checks re-run, with results going somewhere a human reads them.
- Thresholds. Defined levels at which a drift becomes an issue and an issue becomes an escalation. Without these, monitoring produces a report instead of an action.
- An escalation path. A known route from a detected problem to the team that fixes it, with an authority who can reprioritize when that team says no.
Notice what isn’t among those four components: a platform. Tooling makes each of them cheaper and faster to operate, and at enterprise scale across a decentralized estate it’s the difference between a program that runs and one that aspires to. But a platform pointed at a site with no owner produces alerts nobody acts on, which is a more expensive version of the problem you started with. The visibility layer is necessary. It isn’t sufficient, and any vendor telling you otherwise is selling you the easy half.
The transition from project to program is an ownership decision first and a tooling decision second.
So why does continuous monitoring keep beating periodic review? Because quality drifts continuously and reviews don’t. A quarterly audit doesn’t prevent a quarter of accumulated regression, it discovers one. You get a document describing thirteen weeks of degradation, delivered to a team that has already moved on to other work, describing problems whose causes are now hard to reconstruct.
I’ve read those documents. Nobody acts on them.
Continuous monitoring inverts that: small signals, surfaced near the moment they appear, attributable to a specific change.
This is also where the operating model underneath your content function starts to show. Siteimprove’s Content Operations Maturity Model describes the progression from ad hoc operations to governed ones, and a relaunch tends to expose exactly where your organization sits on it. If publishing was ungoverned before the migration, a new platform doesn’t govern it. It just gives the same practices a faster way to scale.
Performance is a standing service level, not a launch benchmark
Performance numbers captured at launch describe a site that stops existing the moment editors start publishing.
That’s not a criticism of launch benchmarking. Baselines are useful, and you need one. But the error most teams make is treating the baseline as an achievement rather than as the starting value of a metric you’ll now watch decline.
Field data is what counts here. Lab scores tell you what your site can do under controlled conditions, and real-user measurement tells you what it actually does across the devices, networks, and pages your audience uses. The Core Web Vitals thresholds give you defensible numbers to govern against: Largest Contentful Paint within 2.5 seconds, Interaction to Next Paint within 200 milliseconds, and Cumulative Layout Shift at or below 0.1, each measured at the 75th percentile of real page loads.
Treat load time, responsiveness, and uptime as service levels rather than scores.
The distinction is operational. A score gets reported. A service level has a threshold, and crossing the threshold triggers something. If your performance reporting produces a monthly number that goes into a deck and nowhere else, you don’t have performance management, you have performance observation.
I’ve sat in those reviews. The number moves, everyone agrees it should move back, and nothing gets assigned.
The question that separates the two is simple: when LCP on your top templates crosses 2.5 seconds, what happens, and who does it?
At enterprise scale that question is harder than it sounds, because the causes of performance drift are distributed across teams that don’t report to each other. Marketing adds a tag. A product team embeds a widget. A regional site publishes uncompressed hero images because nobody told the regional editor about the image pipeline.
Each decision is locally reasonable. The aggregate is a site measurably slower over eighteen months with no single decision you could point to as the cause.
Which means the response path matters more than the alerting. You need a route from a threshold breach to the team that owns the offending component, and an owner who can say no to the thing that caused it. Without that, your monitoring produces a very well-documented decline.
Accessibility and discoverability regress through the same structural inputs
Semantic HTML. Heading hierarchy. Descriptive alt text. Document structure.
Each of those four markup properties sits in your accessibility conformance requirements. Each one also decides whether search engines and answer engines can parse, extract, and represent your content. That overlap is structural and logical: it follows from how both assistive technologies and retrieval systems consume markup, both of which depend on the document telling them what its parts mean.
Let’s be precise about what that does and doesn’t establish. It does not mean accessibility work causes better rankings or higher citation rates, and no public research demonstrates that relationship. What it means is narrower and still useful: when a publishing practice degrades document structure, it degrades both at once, through the same mechanism, in the same release.
An editor pasting styled markup that turns an H2 into bold body text has created an accessibility regression and a discoverability regression in one action.
So monitoring them as two unrelated programs is duplicated effort against a single set of inputs. You run an accessibility audit that flags heading structure. You run a technical SEO crawl that flags heading structure. Two teams, two tools, two reports, two remediation backlogs, but a real chance that neither team knows the other found the same thing. How many times has your organization paid twice to discover one broken template?
Meanwhile the underlying cause, an upload flow or a template or an editor who was never trained, goes unaddressed in both.
The standard against which the accessibility half gets measured is WCAG 2.2, published as a W3C Recommendation in October 2023 and backward-compatible with 2.1 and 2.0, so work you did against earlier versions carries forward. Level AA is what most legislation references.
This is the intersection Siteimprove is built around: connecting accessibility, compliance, content quality, and answer-engine visibility into a single governance layer rather than a set of adjacent tools. It’s a connection that pure-play AEO platforms and traditional SEO platforms don’t make, because neither category treats accessibility as part of the discoverability picture at all. If you’re already investing in accessibility conformance, you’re generating structural properties that serve both purposes. Monitoring them separately means paying for the same information twice and still missing the shared cause.
Feedback without a prioritization rule becomes a queue of complaints
Enterprise teams don’t have a feedback shortage.
You have analytics. You have support tickets. You have accessibility reports, internal publisher complaints, sales anecdotes, session recordings, and an inbox. What you don’t have is a rule for deciding which of that earns engineering time, and in the absence of a rule, the thing that gets fixed is whatever was escalated by the most senior person.
That’s not prioritization. It’s volume control.
I’ve watched a broken contact form sit for two quarters behind a typo somebody’s VP noticed.
A prioritization rule that works ties the fix to two properties you already track: which risk category the issue falls into, and who owns the component. Risk category tells you the exposure, meaning whether this is a compliance problem, a discoverability problem, a content problem, or a nuisance. Ownership tells you whether the fix is a template change that resolves it across the estate or a one-page patch that will recur next month.
Those two properties do most of the work. An accessibility failure in a global template scores high on both and should jump the queue. A broken link on a 2019 press release scores low on both and can wait, or can be resolved by retiring the page.
Neither of those decisions required a meeting, but most organizations hold one anyway.
Regular updates are the other half of this, and they cut both ways. Content and technical updates sustain quality, and they’re also the primary vector by which regressions get reintroduced. Every release is an opportunity for a template to lose a heading level or a new component to ship without keyboard support. An ungoverned update cycle will faithfully reintroduce the exact problems your validation cadence keeps catching, which is how teams end up feeling like they’re fixing the same issue forever.
They are. The fix keeps working. The governance around the release doesn’t.
Launch metrics measure the project, not the estate you now run
Launch metrics answer a question that stops being interesting the day after launch.
Did the migration complete. Did traffic hold. Did the redirects fire. Those are real questions with real answers, but they’re questions about a project. The moment the project ends you’re running an estate, and an estate needs a different measurement model: one that describes state over time rather than the outcome of an event.
Five measures belong in that model, and each one degrades on a different timescale once publishing resumes:
- Accessibility conformance rate, measured as the proportion of your estate meeting your stated WCAG target, tracked as a trend rather than a snapshot.
- Discoverability health, covering crawl coverage, indexation, structured data validity, and the technical signals that determine whether your content can be retrieved and represented.
- Content health, meaning currency, accuracy, duplication, and orphaned pages. Content decay is the slowest of these to show up and the hardest to reverse at scale.
- Performance against thresholds, reported as time spent outside your service levels rather than as an average that hides your worst templates.
- Measurement integrity, which is the check that the other four numbers are real. Broken tracking produces improving metrics.
Those five measures describe the estate. One more describes the program that maintains it, and it’s the one that keeps the program funded.
Measure the program itself.
Mean time from regression introduced to regression detected. Mean time from detected to resolved. Percentage of issues caught by monitoring rather than by complaint. Those three numbers describe whether your governance function is working, but more usefully they’re the only evidence that survives contact with a CFO asking why a line item exists for a site that launched two years ago.
What actually wins that argument? Not a conformance score. An executive doesn’t fund monitoring because conformance is at 94%. They fund it because the last three regressions were caught internally in under a week, and the one before the program existed was caught by a demand letter.
Take Away
Every check in this piece is available to you today. None of them run on their own.
So the last deliverable of your relaunch isn’t the launch. It’s the answer to who owns quality next quarter, and if that question is still open when the project team disbands, you’ve already chosen decay.