How we run the service
Service Level Agreement
DKT Digital, St Peter Port, Guernsey GY1 · info@dktdigital.co.uk
DRAFT — pending legal review, not binding. This document has not been reviewed by a solicitor. It is published so you can read our terms before you talk to us, and so we can be held to what it says — but it is a working draft, not a vetted contract, and nothing here is legal advice. A binding version will be issued for signature only once it has been through legal review. If you are relying on any part of it, ask us and we will tell you where it stands.
How quickly we respond when something breaks, when you can reach us, what monitoring runs without anyone asking, and — stated rather than hidden — what happens if the one person running your automations is unavailable.
Last updated: 2026-08-01. Owner: Daniel Thomas, DKT Digital (Guernsey).
1. What this covers
This SLA describes how quickly DKT Digital responds when something goes wrong with the automations we run for you, when you can reach us, and — honestly — what happens if the person running your automations is unavailable. It is deliberately specific. A vague SLA is worth nothing when a system is down.
It applies to clients on a live monthly retainer (Starter, Growth, or Scale).
2. Support hours
Standard support window: Monday to Friday, 09:00–17:30 UK/Channel Islands time, excluding Guernsey and England & Wales public holidays.
Outside these hours, monitoring still runs automatically (see §5) and critical alerts still reach the owner, but a response is only guaranteed within the window. A P1 issue raised at 18:00 on a Friday has its response clock start at 09:00 the next working day.
3. Priority levels and response times
"Response" means a human has acknowledged the issue and started work — not that it is already fixed. Fix/resolution times depend on the cause and are given as targets, not guarantees.
| Priority | What it means | Examples | Response target | Resolution target |
|---|---|---|---|---|
| P1 — Critical | A live automation that touches money or customers is fully down or doing the wrong thing. | Invoices not being chased at all; onboarding emails not sending; a client-facing automation sending wrong data. | Within 4 working hours | Same or next working day where the fix is within our control |
| P2 — Major | An automation is degraded or partly failing, with a manual workaround available. | One of several automations failing; reports delayed; a non-critical alert channel down. | Within 1 working day | Within 3 working days |
| P3 — Minor | Cosmetic, low-impact, or a change request. | Wording tweak on a report; a new small automation; a question. | Within 3 working days | By agreement |
How to raise an issue: email info@dktdigital.co.uk with the priority in the subject line (for example P1 — invoice chasing stopped). Anything genuinely P1 should also be flagged by the urgent channel agreed with you at onboarding (phone or WhatsApp), because email alone is not guaranteed to be seen out of hours.
4. Escalation path
Buyers treat a missing escalation path as a hidden single point of failure. Here is ours, stated plainly.
- First contact: Daniel Thomas (owner/operator) — email, then the agreed urgent channel.
- If no acknowledgement within the response target: re-send marked "ESCALATION" and use the urgent channel. Automated monitoring will usually have alerted the owner before you do (see §5).
- If the owner is unreachable (see §6): the read-only runbooks come into play, your automations keep running unattended, and your data stays exportable — so you are not left with nothing. A named continuity contact is not yet appointed; §6 is the full and honest account of what that means. This step will name that person once one exists, and not before. (Corrected 2026-08-01: this line previously read "the continuity contact and the read-only runbook come into play", which promised a person who does not exist — in the one section a client reads when something has already gone wrong.)
There is currently one person. §6 is the honest account of what that means and how it is mitigated. We do not pretend there is a 24/7 team.
5. Monitoring — issues are usually caught before you see them
We do not wait for you to report failures. The following run automatically:
- Engine liveness, every 60 seconds: an on-box check that the automation engine is answering, three probes two seconds apart so a single blip does not raise a false alarm. This is our fastest ring. It runs on the server, so it cannot see the server itself failing — that is the next item's job.
- External availability, every 5 minutes: an independent service outside our infrastructure polls the engine's own health path and alerts if it stops answering. Five minutes is the floor of the tier we are on; we say so rather than imply continuous watching.
- Scheduler liveness, every 5 minutes: a check that runs outside the automation engine and reads its heartbeat straight from the database — because a checker living inside the engine dies with it. It waits for three missed beats before speaking, so a stalled scheduler is reported within roughly 15 to 20 minutes. That is deliberate: a shorter window would cry wolf.
- Weekly system report, Mondays: free disk space, memory, container state, whether the tunnel is up, the age of the most recent backup, and whether the database has logged a disk-I/O fault in the previous seven days. This one is weekly and we do not pretend otherwise — it catches slow drift, not outages.
- Each on-box check above verifies that its own alert channel answered before it counts itself as having spoken, and alerts reach a real phone. A monitor whose alerting is broken is worse than no monitor, because it looks fine. External availability is a third party's service and does not report to us on every poll; its alerting was proved end to end on 3 September 2026 from that service's own records.
- Per-workflow error handler: one alert per failing automation per hour (deduplicated, so a storm doesn't bury the signal), sent to the owner immediately.
- Free-tier quota monitor (daily): warns before an email/verification quota is hit, so sends don't silently start failing.
Monitoring that has never been tested against a real outage is only a hypothesis, so we run documented incident fire-drills (see our internal incident runbook) and publish what they measure — including when the answer is unflattering.
The last drill was 2 September 2026, on the current server. The automation engine was deliberately stopped for 41 seconds. Nothing detected it. Our external monitor polls every five minutes and the outage fell between two polls; the scheduler check is designed to wait fifteen minutes before speaking. Both behaved exactly as built, and neither was built for that question.
So we built the missing one. A second drill the same day, after adding a 60-second on-box liveness check, measured detection at 39 seconds. On 3 September we then proved the alerting end to end — a deliberate failure raised an alert, the fix cleared it, and both were confirmed from the monitoring service's own records rather than from its dashboard, with **no interruption to the live service**.
What that does and does not mean. The 39-second detection is from a check that runs on the server. It is fast, and it is blind to the server itself failing, to the network tunnel dropping, or to a fault at our CDN. Those are the external monitor's job, and it polls every five minutes — the floor of the tier we are on. **A very short total outage may therefore still pass unseen from outside.** We would rather write that down than let "monitored" stand for something faster than what runs.
6. Continuity — the honest part
**DKT Digital is currently a one-person business. Daniel Thomas is a single point of failure.** Hiding that would fail any serious buyer's diligence; stating it with a mitigation is the honest and stronger position.
What keeps running with no human present:
- Every live automation (they run on a schedule or on triggers, unattended).
- All monitoring and alerting listed in §5.
- Nightly backups, replicated off-site to encrypted cloud storage (Security Overview §5).
- The client portal and website.
So if the owner is unreachable for, say, a week, your automations keep working. What does not happen without a human is: changes, new automations, investigation of a novel failure, and responses to support requests.
What happens if the owner is unreachable for an extended period:
- Continuity contact (to be arranged — a named trusted technical contact): holds sealed access instructions and can, at minimum, stop or pause automations and notify clients. This is not yet in place and is on the critical path — see the note below.
- Read-only runbook (our internal incident runbook, our internal restore runbook): documents how the system is structured, how to restore from backup, and how to reach clients.
- Your data is portable at all times (see Offboarding): even in a worst case, you can take a complete export of your data in standard formats and move to another provider. You are never locked in.
Where this stands today, stated plainly: the continuity contact is not yet appointed. Until it is, the honest position is that your automations keep running unattended, your monitoring keeps running, and your data stays fully exportable at any time — but there is no second person who can make changes on your behalf. We would rather tell you that than let you discover it. This SLA should not be signed as promising a continuity contact until one exists.
7. Service credits (proposal — confirm with solicitor)
If we miss a P1 response target in a given month, a proposed remedy is a service credit of 10% of that month's fee per missed P1 response, capped at one month's fee. This aligns incentives without exposing a small business to uncapped liability. Final wording, caps, and exclusions (e.g. third-party outages outside our control — see §8) are for the solicitor.
8. What is outside this SLA
Response/resolution targets do not apply where the cause is outside our control, including:
- Outages at a third-party sub-processor (e.g. the client's own Xero, a payment provider, an email provider) — we will chase and work around, but cannot guarantee their timing. The full sub-processor list is in Security Overview.
- The client not providing access, credentials, or information needed to resolve an issue.
- Force majeure.
- Free-trial or discovery-audit engagements (no retainer = no SLA).
*Related: Security Overview · Offboarding · Scope & Inclusions · our internal incident runbook · our how-it-works notes.*