Summary
This page publishes what automated probes observe from the Cloudflare network. No value is entered by hand, no history is reconstructed after the fact, no percentage is rounded.
The document below describes the measurement precisely enough that you can dispute it. If you observe something that contradicts what is displayed, the procedure is at the end.
Monitoring active since August 3, 2026.
Where the measurements originate
The probes run inside a Cloudflare Worker triggered by a scheduled task. They therefore originate from the Cloudflare network, not from a workstation, not from a server at Webryk, not from your office.
Three consequences, to keep in mind before reading a single number:
- The exit point is not pinned to a specific data centre. Cloudflare decides where the task runs. We do not publish the execution location because we do not control it.
- The network path between a Cloudflare point of presence and an origin is not the path between your internet provider and that same origin. “Operational” here means “reachable from the Cloudflare network”, not “reachable from where you are”. A peering failure, a corporate DNS resolver, a local firewall or a mobile carrier can cut you off from a service this page reports as operational. Both statements can be true at the same time.
- Requests come from Cloudflare address space. Some third-party vendors treat that traffic differently from residential traffic, particularly for rate limiting and bot detection.
There is no real-browser measurement. This page collects nothing from its visitors.
What is measured, and how often
| Check type | Frequency | In service | Scope |
|---|---|---|---|
| Black-box HTTP | 60 s | yes | Public surfaces: marketing site, client portal, partner area, document vault, public demo |
| Application health endpoint | 60 s | yes | Structured response from a dedicated endpoint reporting the state of each capability’s internal dependencies |
| Vendor status feeds | 5 min | yes | Public status pages and feeds of the third-party vendors listed |
| Synthetic sign-in journey | 15 min | no | Full sign-in with second factor, through to an authenticated page |
| DNS resolution | 60 min | yes | The A record of the monitored domain and the MX records for email |
| TLS expiry | 60 min | no | Days remaining on the certificate and integrity of the chain |
| Scheduled-job heartbeat | 60 min | no | The absence of a signal emitted by a daily job when it finishes |
Three of these six instruments are not in service today. The reasons are in the matching sections below, and the consequence is the same in all three cases: where an instrument is missing, nothing is measured, and nothing measured means “unknown”, never “operational”.
Four rows are in that position: “Sign-in and MFA”, “Payments”, “Scheduled jobs” and “Stripe”. None of them publishes a percentage, and the figure slot says so in plain words rather than announcing a count that would never move. A real check that fails on one of them is still displayed as it stands: measurement always beats a declaration of absence, and “monitoring not yet in service” will never be shown over an observed outage.
Scheduled jobs are a special case: they are not probed, they report themselves. See below.
Black-box HTTP check (60 s)
GET request from the Worker, or HEAD on the targets where the headers are enough, no caching, abandoned at 8 seconds. Past that the request is dropped and the measurement counts as a failure, not as an absent measurement: the target had its time and did not answer.
Redirects are not followed. A redirect is a signal about the target, not a detour to be taken, and it is judged as it stands: a surface that was to answer 200 and answers 301 is not conforming. Some targets accept a redirect as a conforming answer because that is their normal behaviour; that is declared target by target, never by a blanket rule.
The check passes if the response code is the one expected for that target and if the response body contains the expected marker where a marker is defined for that target. A 200 served by a generic error page does not count as a pass. The connection is HTTPS and the runtime refuses an invalid certificate, which fails the request; that does not mean the certificate is measured, and the TLS section below says why.
Total response time is recorded. It does not by itself trigger an outage, and there is no single threshold: each check carries its own. Public HTTP surfaces move to “degraded” above 2,000 ms, application health endpoints above 1,500 ms, and the observed GitHub probe above 3,000 ms. A single threshold announced for all of them would be false for two families out of three.
For this type of check, the published rule is that a single failed measurement does not change the displayed state: three consecutive failures are required, which is three minutes at this cadence, and returning to “operational” requires as many. That threshold of three applies to the 60 s checks, and to them only: TLS, DNS and scheduled-job heartbeats change state on a single observation. Both rules are stated in full in the next section.
That damping is not applied to the display yet, and saying so beats letting you assume a smoothing that does not exist. Today a row presents the most recent verdict of each of its checks: a single failed measurement can therefore turn it red, and a single conforming measurement can turn it back. The thresholds above are declared in advance and published here so that they are checkable on the day they govern the display. The gap runs toward sensitivity, not toward flattery: the current display can show an outage that damping would have absorbed, it cannot hide an outage that damping would have held back.
When a state change is published, the start timestamp will be that of the first failed check, not the third. We will not push an outage’s start forward to make it shorter. No state-change instant is published today.
Application health endpoint (60 s)
Each application capability exposes a health endpoint that returns a structured state of its dependencies. It distinguishes “the page loads” from “the page loads but the function behind it does not work”.
The content of that response is never published. No internal service name, no error message, no diagnostic code, no database latency. That information is worthless to a visitor and of definite value to an attacker. Only the consolidated public state of the row is displayed.
Vendor status feeds (5 min)
We read each vendor’s public status page or feed every 5 minutes and map its declared severity onto our five states. A document older than 30 minutes yields no verdict rather than carrying its last value forward: a frozen feed is not an operational vendor, it is an unreadable feed.
The lag is real and we do not correct for it: on top of our 5-minute polling comes the time the vendor takes to publish, often 5 to 30 minutes, sometimes several hours, sometimes never for a partial incident. What we show in the vendor column is therefore what the vendor was saying at the moment we read it, not what was actually happening.
We take a snapshot at each read. If a vendor later edits or deletes its own history, our snapshot stays as it was.
Synthetic sign-in journey (15 min): not in service
This journey does not run today. It needs a dedicated canary account holding the narrowest possible rights, and that account does not exist yet. Until it does, nothing is measured on the rows it governs, and this page publishes neither a concluded state nor a percentage for them. What follows therefore describes the instrument as declared, not a measurement under way.
As declared: a dedicated test account, created for this purpose alone, runs the full four-step sequence every 15 minutes: load the sign-in page, submit credentials, resolve the second factor, reach an authenticated page. A total budget of 25 s for the sequence, with a hard abort at 30 s. Failure at any step fails the journey.
That account will belong to no client, will contain no client data and will have access to no client data.
A direct consequence worth holding on to, for the day it runs: the “Sign-in and MFA” row will have a resolution of 15 minutes. A 10-minute sign-in outage can pass entirely between two runs and leave no trace. Since two consecutive failed runs are required to publish a change, which is 30 minutes at this cadence, a sign-in outage may only appear on this page 30 minutes after it began.
DNS resolution (60 min)
Every hour, over DNS on HTTPS: resolution of the A record of the monitored domain and of the MX records for email, checking the answers against the expected records. Abandoned at 10 s.
A single observation is enough to publish a state change: a DNS record is not a flaky signal, and requiring a second measurement would cost a full hour and prove nothing. A failure to resolve is an outage happening now, and it is treated as one.
TLS expiry (60 min): not measurable, and the published state says so
This check cannot answer the question it is asked, and records “unknown” on every run. Nothing inside a Cloudflare Worker observes the certificate of an outbound request: a fetch() response exposes no peer certificate, and the only pair of dates the runtime documents describes the incoming client’s certificate on a mutually authenticated connection, which is a different thing entirely. There is therefore no measurement today of the days remaining on a certificate or of the integrity of a chain.
Recording a pass because a connection opened would be publishing an expiry measurement nobody took. Every run therefore records “unknown” with its reason, honestly and by construction. A failed handshake also records “unknown” rather than “down”, for the same discipline in the other direction: the platform returns one undifferentiated error for a DNS failure, a refused connection and a rejected certificate alike, so this probe cannot conclude that the certificate is what failed.
What this does not cost: a genuinely expired certificate is already an outage that the 60-second HTTP check reports, at 60-second resolution, because the platform refuses the handshake. What it does cost: the advance warning. The silence of this instrument is set aside from the arithmetic of the rows it accompanies, because “nobody measured the certificate” is not an answer to “is the site answering”. It is the only check in that position, and it is the only one because it is the only one asking about a future risk rather than about the present.
Scheduled jobs (daily heartbeats): not in service
No heartbeat can arrive today, because the door is not open. The receiving route is written and shipped inside the Worker, and it is not published: reaching it needs a domain the deploy token has no right to create. “No heartbeat has ever arrived” and “no heartbeat can arrive” are two different admissions, and it is the second one that is true. The “Scheduled jobs” row therefore measures nothing.
As declared: three daily jobs emit a heartbeat when they finish. The measurement is of a missing heartbeat, not of an outbound request, and it is evaluated every hour.
The deadline is counted from the last heartbeat received, never from a schedule: 24 hours of expected interval, plus 2 hours of grace. Past that the row goes down on a single observation, because a missed daily job is not noise. That grace is deliberately wide, and what it changes is worth saying rather than discovering: this instrument detects a job that did not run, it does not detect a job that ran late. A tighter instrument would publish outages that are not outages until the scheduler’s precision has been measured, and a false outage costs more here than an hour of ignorance about a daily job. The identifiers, exact schedules and names of these jobs are not published.
A failed scheduled job affects no interactive surface. The two rows are independent and will stay that way.
Publication thresholds: two rules, not one
How many non-conforming observations does it take before a state change is published? The answer depends on the type of check. There is no single blanket rule, and this page no longer claims one.
| Check type | Cadence | Failures before publication | Sustained failure | Successes before recovery | Sustained recovery |
|---|---|---|---|---|---|
| Black-box HTTP | 60 s | 3 | 3 min | 3 | 3 min |
| Application health endpoint | 60 s | 3 | 3 min | 3 | 3 min |
| Vendor status feeds | 5 min | 2 | 10 min | 2 | 10 min |
| Synthetic sign-in journey | 15 min | 2 | 30 min | 2 | 30 min |
| Scheduled-job heartbeat | 60 min | 1 | 60 min | 1 | 60 min |
| TLS expiry | 60 min | 1 | 60 min | 1 | 60 min |
| DNS resolution | 60 min | 1 | 60 min | 1 | 60 min |
The two rules, stated:
- Fast checks, at a 60 s cadence (HTTP and health endpoints), require three consecutive failures before any published state change. At that cadence it is three minutes of agreement between three measurements: long enough that a single lost packet cannot turn a row red, short enough that a real outage is on the page before a client has finished writing the email reporting it.
- Slow checks (TLS, DNS, scheduled-job heartbeats) change on a single observation. These are not flaky signals. A certificate expiry date does not flicker, neither does a DNS record, and a missed daily job is not noise. Waiting for a second observation would cost 60 minutes and would prove nothing.
- Vendor feeds and the synthetic journey sit between the two, at two observations, which is 10 and 30 minutes respectively.
The durations in the fourth and sixth columns are not typed in by hand: they are the cadence multiplied by the number of measurements required. That is the only honest way to state them.
The three instruments that are not in service appear in this table because their thresholds are declared in advance. While they measure nothing they publish no state change, and none of these durations ever runs.
The number of successes required to publish a return to “operational” equals the number of failures required to publish an outage. An asymmetric pair would invite the question of which direction Webryk is being generous in; symmetry answers it: neither.
The five states
State is never signalled by colour alone. Every state carries a distinct shape and a written label, including inside the 90-day bar, and stays readable in greyscale.
| State | Shape | What triggers it |
|---|---|---|
| Operational | Disc | Every applicable probe returned a verdict, and that verdict was conforming: expected response, within the expected time budget, with valid content. |
| Degraded | Triangle | The capability responds, but outside thresholds or only partially. Three cases: response time above the threshold belonging to that check, on three consecutive checks at the 60 s cadence (2,000 ms on public HTTP surfaces, 1,500 ms on health endpoints); a health endpoint reporting a failed dependency while the surface still responds; only some of the targets in a single row failing. Degraded time does not count as uptime. |
| Down | Octagon | The number of failed observations required by that type of check has been reached: three for the 60 s checks, two for vendor feeds and the synthetic journey, one for TLS, DNS and heartbeats. No response, timeout, 5xx code, invalid TLS, or failure of the synthetic journey at any step. |
| Planned maintenance | Diamond | A period inside a window announced at least 48 hours in advance. Measurement continues during maintenance and the observed results are recorded, not erased. |
| Unknown | Ring | No verdict could be reached for the interval: the probe itself failed, the vendor feed was unreadable, or no measurement exists for that period. |
An unannounced intervention, or one announced with less than 48 hours’ notice, is not maintenance. It is counted as down. Otherwise, announcing late would be enough to erase an interruption.
The overall state shown at the top of the page reflects Webryk capabilities only. Third-party vendor rows never change it, in either direction.
Resolution: what the interval cannot see
An interruption shorter than the check interval can pass entirely between two measurements and leave absolutely no trace. The shortest outage this instrument can see lasts 60 seconds, because that is the fastest cadence it runs.
Concretely: a 40-second outage on a surface checked every 60 seconds is probably invisible here. A 10-minute sign-in outage on a journey that runs every 15 minutes is probably invisible here. That does not mean it did not happen, it means this instrument does not see it.
Restarts, fast failovers and most intermittent errors affecting a fraction of requests fall into that blind spot. A sampling probe does not measure an error rate: it measures instants. If 2 % of requests fail continuously, there is roughly a 2 % chance the probe lands on one at any given instant, and the row will very probably stay “operational”.
This is not a defect to be fixed in a later version, it is the structural limit of the method. You need to account for it when interpreting anything shown here.
Unknown periods
An interval is marked “unknown” when we have no verdict we can trust: an instrument not yet in service, Worker execution failed, write to history failed, or a vendor feed unreachable or unreadable.
Any period before monitoring began is unknown for the same reason: nobody was watching.
Monitoring began on August 3, 2026.
Unmeasured time is always unknown. There is no tolerance, no margin and no setting that can soften it: a sample speaks for its own cadence, from the instant it was taken, and not one millisecond longer in either direction. Painting a measurement gap with the next sample’s verdict would be backfill, and it is banned.
The arithmetic applied:
- Unknown minutes are excluded from both the numerator and the denominator. They count neither as available nor as unavailable.
- The total unknown minutes in the window is published next to the percentage. It is not hidden in a footnote.
- If unknown minutes exceed 5 % of the window, the percentage is not published at all and the reason is displayed in its place, with the measured share: “Percentage withheld - P% of the window unmeasured, above the 5% ceiling”. P is the real share of unknown time, truncated to two decimals like every other percentage on this page. Only the raw counts remain displayed.
That 5 % ceiling applies to the whole window. It is not the same rule as the 90 % daily coverage floor, which only decides whether a day counts toward the 30-day gate. Thirty days each scraping past 90 % coverage would clear the daily threshold while a tenth of the assembled window had never been observed, and a percentage computed from that would be a percentage about a period Webryk did not watch.
Treating unknown time as available time would be the flattering choice. It is also the choice that makes a number unverifiable.
Uptime: two figures, not one
Rolling 90-day window, minute granularity. Two figures are published side by side:
- Uptime including planned maintenance. Every minute during which the capability did not deliver the expected service counts as unavailable, announced maintenance included.
- Uptime excluding planned maintenance. Minutes falling inside a window announced at least 48 hours in advance are removed from the denominator.
What each state becomes in the calculation, with no exception and no option:
| Time spent in state | In the numerator | In the denominator, maintenance included | In the denominator, maintenance excluded |
|---|---|---|---|
| Operational | yes | yes | yes |
| Degraded | no | yes | yes |
| Down | no | yes | yes |
| Planned maintenance | no | yes | no |
| Unknown | no | no | no |
Degraded time is excluded from the numerator. Said in plain words, because an undocumented exclusion is the same class of defect as a wrong number: a capability responding outside its thresholds is not an available capability, so those minutes are not counted as uptime. They do stay in the denominator: they are not erased, they are counted as time during which the service was not delivered as expected. There is no setting that decides otherwise, because the only function such a setting could have is to make a bad month look better.
Publishing only the second figure would be misleading, for two reasons. First: the person who could not sign in at 2 a.m. on a Sunday could not sign in, announced or not; the announcement changes the courtesy, not the availability. Second: if maintenance leaves the denominator, it becomes possible to improve the figure by scheduling more interruptions. Any measure you can improve without improving the service stops being a measure.
Publishing only the first would be unfair to maintenance done properly. Hence both.
Percentages are truncated toward zero at the second decimal, never rounded. Two worked examples: uptime measured at 99.994 % displays as 99.99 %; uptime measured at 99.999 % also displays as 99.99 %, never 100.00 %. Rounding is the only arithmetic operation on this page capable of making a measurement look better than it was, so it is not available anywhere. “100 %” is never displayed if any interruption was recorded in the window.
Why the percentage is withheld for 30 days
Until 30 days of real data exist, the percentage slot reads “Collecting data - N of 30 days”, where N is the number of qualifying days. No percentage is published before that.
The reason is arithmetic. The same 10-minute interruption is worth 0.23 % over a 3-day window (4,320 minutes) and 0.007 % over a 90-day window (129,600 minutes). The same incident would therefore produce “99.76 %” or “99.99 %” purely as a function of window length. A percentage computed over a few days measures the window, not the service.
There is also a starting bias: a short, recent window is very likely to be clean, which produces a beautiful number that says nothing. Publishing “100 %” after four days would be strictly true and usefully dishonest.
The two withholdings are evaluated in this order: first the 30 qualifying days, then the 5 % unknown-time ceiling described above. The order matters. A window in its first week reads “Collecting data”, not “Percentage withheld”.
The qualifying-day counter itself is visible from the first measured day. There is nothing to hide in having started.
Before that there is no counter, and that is deliberate. “Collecting data - 0 of 30 days” would assert that collection is under way, which is a claim about a system nobody has observed running. Until a probe has reached a verdict, the slot says that nothing has been measured, and it says it without a figure. A row whose instrument is not in service does not show that counter either, ever: on it the count would stay at zero forever, and a figure that cannot move is a promise nobody is keeping.
Third-party vendors: two independent signals
Each vendor row carries two values that are never merged:
- Vendor reports: what its public status page showed at our last read, at most 5 minutes earlier.
- Webryk observes: what our own probes find on the vendor endpoint that Webryk actually uses.
These two signals diverge, in both directions. A vendor can report itself operational while we see failures, because the outage is regional, partial, limited to one specific API, or simply not published yet. It can also report an incident while everything works for us, because the incident affects a region or a product we do not use.
When the two signals diverge, the row is flagged “signals diverge” and both values stay displayed. We do not pick whichever version suits us, and we do not overwrite the vendor’s statement with our observation. The divergence is itself the most useful information in the row.
We have no privileged access to these vendors’ infrastructure. Our observation is worth exactly what a black-box measurement taken from the outside is worth.
Incidents
Automatic opening is not yet in service. No component in production opens, updates or closes an incident today. Any incident published before that is written by a person and published by a rebuild of the site. The rules below are the ones that will apply, and they are written here in advance so that they are verifiable from the outside on the day they do.
An incident will open automatically when an outage is confirmed, that is, when the number of failed observations required by that type of check has been reached. The system will then publish a single pre-written sentence, identical in French and English, saying only what has been observed.
Everything after that first sentence is written by a person. No incident update is generated automatically.
An incident will close automatically after sustained recovery. That duration is not the same everywhere: it is the cadence of the check multiplied by the number of conforming measurements required, which is 3 minutes for the 60 s checks, 10 minutes for vendor feeds, 30 minutes for the sign-in journey, and 60 minutes for TLS, DNS and heartbeats. A single figure announced for all seven types would be false for six of them. An incident cannot close less than 5 minutes after it opened in any case, so that a flapping service does not produce a series of short incidents that would hide how bad the hour actually was.
A run of unknown periods never opens a public incident. Not seeing is not the same as observing an outage, and publishing one in place of the other would be an invention. After five consecutive unknown evaluations, an internal alert goes to operations.
Automatic closure closes the state, not the file: post-incident analyses are always written by a person and published separately.
An incident’s timestamps are those of the measurements, not those of the writing.
The infrastructure behind this page
This page is deliberately built on a stack that does not overlap with the one it monitors.
- Static site built ahead of time, served by Cloudflare Pages.
- Probes running in a Cloudflare Worker on a scheduled task.
- History stored in Cloudflare D1.
It uses no Vercel, no Supabase, no Resend and no component of the Webryk application platform. No call to the portal database, no authentication, no application API.
The precise failure it is designed to survive: a total outage of Vercel, Supabase or Resend makes Webryk services unusable and leaves this page online, current, and able to say that those services went down. That is exactly the moment a status page is worth anything, and exactly the moment a status page hosted on the same platform as the product disappears along with it.
Two residual dependencies, stated plainly:
- A global Cloudflare outage takes this page down with it. There is no second host.
status.webryk.cashares the DNS zone ofwebryk.ca. A failure at the registrar or zone level therefore hits both at once.
We do not claim to be independent of everything. We are independent of what we measure.
What this page does not measure
Read this before concluding anything from an operational row.
- Actual email delivery. We check reachability of the sending API, resolution of the MX records and the vendor status feeds. We do not send an end-to-end test message to a third-party inbox on every cycle. An email can therefore leave correctly and land in a junk folder, or be held by the recipient, without anything appearing here.
- Actual payments. We check reachability of the payment path and the state declared by the vendor. We run no test transactions. A decline caused by an issuing bank or an antifraud rule is not visible here.
- Your experience from your network. See above. This is the most frequently misunderstood limit.
- Perceived performance. We record a server response time, not a browser rendering time, not a complete user journey, not field data.
- Data correctness. An application that responds quickly and shows wrong data is “operational” as this page uses the word.
- Security posture. A status page is not a security dashboard.
- Client-owned systems. Webryk also monitors systems it does not own. They are not published here, because their state belongs to their owner, not to us. The published inventory covers every capability Webryk operates: no Webryk capability is removed from this page because it is behaving badly.
What these numbers are not
They are measurements observed by automated probes. They are not a service level commitment, not a warranty, not a contractual term, and not a promise about the future.
This page describes what happened. It does not describe what is promised. Contractual commitments, where they exist, are in the signed contract and are independent of what is shown here: nothing on this page creates, extends or restricts them.
We say this plainly rather than defensively, because it is the only honest reading of a black-box measurement carried out by the party being measured.
Retention and corrections
What is kept, and for how long:
| What is kept | Duration |
|---|---|
| Raw results of every check, to the second | 14 days |
| Snapshots of the vendor status feeds | 14 days |
| Daily summaries per capability, which the 90-day bar is drawn from | indefinitely |
| Incidents and post-incident analyses | indefinitely |
The published window is 90 days; raw retention is 14 days. These are not two versions of the same number. Two weeks of full-resolution measurements is enough to analyse an incident, and beyond that the storage would cost six times more for forensic value that collapses. What feeds the 90-day window after the prune is the daily summaries, which are never deleted and whose sealed days are protected against rewriting by the database itself.
One consequence worth stating rather than leaving to be discovered: a day older than the 14-day horizon can no longer be attested from the raw measurements, so it cannot be counted as a qualifying day by that route. A percentage nobody can substantiate is not published.
History is not rewritten. No data is added retroactively to fill an unknown period, nor to fill a period before monitoring began.
If a published figure turns out to be wrong, the correction is appended and dated, and the original value stays visible along with a note on what changed. A silent correction would be indistinguishable from falsification.
To dispute a measurement, write to contact@webryk.ca with: the timestamp in UTC, the capability concerned, what you observed, and the network you were on. If our record is wrong, we say so on this page.
The correction we did not make: GitHub, 3 August 2026
On 3 August 2026, seven checks were recorded as “down” for the GitHub row. GitHub was not down. Those seven measurements are wrong.
The cause: the 60-second check targeted api.github.com, an endpoint whose quota is 60 requests per hour per source IP address, and Cloudflare Workers egress from a shared, rotating pool of addresses. The probe therefore spent a quota strangers had already partly consumed, received an HTTP 403 rate-limit response, and classified it as a failure of the service. At the same moment, github.com answered 200 from the same machine. The rule that should have recognised a rate limit did not fire, because it required a header GitHub does not send in that case. That check now targets github.com.
The seven rows stay exactly as written. We do not correct, reclassify or delete a measurement recorded in production, including when we are certain its conclusion is wrong. Correcting the state to “unknown” would have been defensible; deleting the rows would not have been. Both options were rejected for what they would establish: the first time the record is touched to make a bar look better, its immutability stops being a property and becomes a preference.
What that costs, unsoftened: the GitHub row carries a visibly wrong red segment for 3 August 2026, and the uptime arithmetic for that row counts seven checks of downtime that did not happen. That day’s bar is wrong and will stay wrong until it ages out of the 90-day window. Those seven rows will reach no published percentage: 3 August 2026 is the first day of a 30-day gate, and they leave the window before any figure can be published for that row.
The decision and its reasons are recorded in architecture decision 0004, “The record is not edited, even when we are certain it is wrong” (docs/adr/0004-the-record-is-not-edited.md).