Skip to content

Database — Non-Functional Requirements & Our Approach (Tender Response Library)

Reusable across all tenders. A structured catalogue of the non-functional requirements a tender states for a database, each paired with how we deliver it and the measurable evidence we commit to. No commands — this is the approach document (the engineering “how” lives in [[PostgreSQL-HA-Production-Install-Guide]]).

How to use this document

  1. Find the NFR the tender asks for (organised by the ISO/IEC 25010 quality model + the common tender headings).
  2. Copy the Requirement / Our approach / Evidence & target block; adapt the target to what the tender demands and what the awarded architecture actually supports.
  3. Keep every answer honest and measurable — we win on evidence, not adjectives.

The answer pattern (use it every time). Three moves:

① Restate the requirement in our words (shows we understood it) → ② State the concrete mechanism (the architecture/process that meets it) → ③ State the measurable evidence (the number we commit to and how we prove it).

Reference solution these answers assume. A 3-node PostgreSQL cluster with Patroni automatic failover, an odd-numbered etcd consensus quorum, dual HAProxy + Keepalived virtual IP, PgBouncer pooling, and pgBackRest backups with continuous WAL archiving and point-in-time recovery — an all-open-source, standards-based stack with no proprietary lock-in. Substitute your committed versions/targets per tender.

⚠️ Golden rules for NFR answers

  • Never over-bid the SLA. A 3-node cluster honestly underwrites 99.95–99.99 % at the database tier. Reflexively promising “five nines” (99.999 %) is a red flag to a technical evaluator and a trap in the contract.
  • Scope the SLA to the database endpoint, with explicit maintenance-window and upstream-dependency exclusions.
  • Never claim RPO = 0 while running asynchronous replication. State the real bounded RPO instead.
  • Differentiate on evidence — drill reports, SLO dashboards, restore logs, benchmark scans — not on marketing words.

  • Requirement. The database service shall provide a defined availability level, measured at the client-facing endpoint, excluding pre-agreed maintenance.
  • Our approach. A three-node PostgreSQL cluster under Patroni with automatic leader election and an odd-numbered etcd quorum; clients reach a single virtual IP through dual HAProxy + Keepalived, so a node failure is transparent. Planned maintenance is done by controlled switchover so patching does not consume the availability budget. Availability is measured independently of any load-balancer health check, using a synthetic PostgreSQL-protocol probe.
  • Evidence & target. 99.95 % monthly (≤ ~21.9 min/month; a 3-node Patroni cluster realistically underwrites 99.95–99.99 %), measured as successful synthetic probes over a rolling 28-day window on a Prometheus SLO dashboard with error-budget burn-rate alerting. State one measurement window and one number — do not mix “monthly” and “28-day rolling”.
  • Requirement. Loss of the primary shall be detected and a replica promoted automatically, with no operator intervention, within a defined recovery time.
  • Our approach. Patroni holds a leader lock in etcd; on primary loss the lock expires and the most-current eligible replica is promoted, while HAProxy re-routes the write endpoint via a Layer-7 health check against Patroni’s REST API. Detection/promotion timing is tuned to the network to hit the target without false failovers.
  • Evidence & target. We commit to three distinct recovery numbers (they are different and a serious evaluator expects all three):
    • Automatic-failover RTO ≤ 30 s (typically 10–25 s) — Patroni promotion.
    • End-to-end client-recovery RTO — promotion + load-balancer detection + pool reconnect (a small number of additional seconds).
    • DR-restore RTO — full rebuild from backup after total cluster loss (minutes to hours, scaling with database size + WAL replay). All validated by timed failover and restore drills with captured timestamps.

1.3 Fault tolerance / no single point of failure

Section titled “1.3 Fault tolerance / no single point of failure”
  • Requirement. The clustering layer shall tolerate loss of at least one node with no loss of service; the topology shall have no single point of failure.
  • Our approach. An odd-sized consensus quorum (3 tolerates 1, 5 tolerates 2 — never 2, which tolerates zero); three database nodes; dual load balancers behind a virtual IP whose failover is driven by a weighted liveness probe of the load-balancer process; node-level fencing (watchdog). No consensus, database or load-balancer component is singular.
  • Evidence & target. Documented fault-tolerance matrix (N members tolerate ⌊(N-1)/2⌋ losses); a witnessed test showing the quorum stays healthy with one member down and the virtual IP migrating when a load balancer is killed.

  • Requirement. The database shall bound or eliminate committed-transaction loss on failover.
  • Our approach. The durability posture is an explicit decision made with the client, because it is a genuine trade-off:
    • Zero-loss (RPO = 0): Patroni-managed synchronous replication — a commit is acknowledged only after at least one standby has persisted it, and only a synchronous standby is ever promoted.
    • Latency-optimised: asynchronous replication with a bounded lag ceiling, giving a small, stated RPO.
  • Evidence & target. For zero-loss: RPO = 0, proven by kill-primary tests showing the promoted node holds the last acknowledged commit. For the async posture: a bounded RPO stated in real seconds (never “0”). ⚠️ Reconciliation with availability: strict zero-loss mode refuses writes if no synchronous standby is available, which can reduce write availability. We resolve this explicitly per tender by either (a) sizing enough synchronous candidates that a single loss cannot stall writes (quorum-based synchronous commit), or (b) an agreed SLA carve-out that a synchronous- standby-loss write-stall does not count against the availability budget. We state which, up front.
  • Requirement. The solution shall support recovery to any point within a retention window, with off-host, off-site, immutable, encrypted backups.
  • Our approach. Physical backups (full + differential + incremental) with continuous WAL archiving enabling point-in-time recovery, taken from a standby to offload the primary, stored in encrypted object storage with immutability (write-once/object-lock) so backups survive ransomware or accidental deletion. Backup is treated as a distinct control from replication — replication faithfully propagates a DROP TABLE or corruption to every standby; only PITR recovers from it.
  • Evidence & target. PITR to an arbitrary time within retention; 3-2-1-1-0 satisfied (≥3 copies, 2 media, 1 off-site, 1 immutable, 0 restore errors); WAL archive gap count = 0; retention e.g. 35 days online; RPO for the backup path bounded by the WAL archive timeout (e.g. ≤ 5 min).
  • Requirement. Recoverability shall be demonstrated by regular disaster-recovery testing, not by backup-success flags alone.
  • Our approach. Scheduled automated restores into an isolated environment with application-level consistency and row-count checks, plus rehearsed single-node and full-cluster rebuilds; every drill records the achieved RPO/RTO and the measured restore duration for the real database size.
  • Evidence & target. Monthly successful restore test + quarterly timed end-to-end DR drill, each producing a dated report showing restore-to-serving within the committed RPO/RTO.
  • Requirement. The stored data shall have a high durability guarantee.
  • Our approach. Synchronous WAL persistence across nodes, page-level checksums, full_page_writes, and independently-retained immutable backups mean data survives node, disk and site-level events.
  • Evidence & target. Stated annual durability objective backed by the replication + immutable backup design; checksum verification and pgBackRest verify runs with zero errors.

  • Requirement. Authentication shall use a salted, cryptographically strong mechanism; weak/legacy hashing (MD5) is prohibited.
  • Our approach. Cluster-wide scram-sha-256, enforced on every host-based-authentication rule, with channel binding to block downgrade/man-in-the-middle; certificate authentication available for high-value service accounts.
  • Evidence & target. Zero roles with legacy (MD5) password verifiers; 100 % of remote authentication rules use scram-sha-256; no trust/password rules.
  • Requirement. All data in transit, including replication, shall be encrypted (TLS 1.2+).
  • Our approach. TLS enforced through SSL-only host rules (not merely “TLS enabled”), modern ciphers, clients validating the server certificate fully; replication traffic encrypted, optionally with mutual TLS.
  • Evidence & target. No non-TLS remote rules; minimum protocol ≥ TLS 1.2 (target 1.3); TLS confirmed active on all non-loopback backends including the replication stream.
  • Requirement. Personal/sensitive data shall be encrypted at rest across its full lifecycle.
  • Our approach. Volume-level encryption on the data, WAL and backup storage (open-source PostgreSQL has no built-in transparent data encryption, so this is done at the block layer), keys held in a KMS/TPM separately from the database administrators; column-level encryption for the most sensitive fields; backups encrypted independently.
  • Evidence & target. Encrypted block devices confirmed for data, WAL and backup mounts; documented key management owned separately from DBAs; coverage extends to replicas, temp and backups.
  • Requirement. Access shall follow least privilege; no shared or superuser accounts for applications; every actor individually attributable.
  • Our approach. Per-service login roles with only the grants they need (no superuser, no role/DB-creation), separated from object-owner roles; default-privilege and public-schema lockdown; row-level security for multi-tenant data; superuser reserved for named break-glass administrators.
  • Evidence & target. Application roles carry no superuser/create privileges; CREATE on the public schema revoked from PUBLIC; superuser count minimised and each maps to a named individual; grants reviewed and documented.
  • Requirement. No credential shall be stored in plaintext; secrets shall be centrally managed, rotated and auditable.
  • Our approach. A secrets manager (e.g. HashiCorp Vault) issues dynamic, rotated database credentials delivered as tightly-permissioned files — never in environment variables, config maps or committed manifests — with issuance auditing.
  • Evidence & target. Zero plaintext database passwords in repositories/config/env (secret-scan clean); an audit trail of credential issuance; a documented rotation interval.
  • Requirement. Security-relevant database activity shall be logged, retained, tamper-resistant and reviewable.
  • Our approach. Scoped audit logging (schema changes, privilege changes, writes, plus object-level logging on regulated tables — never blanket “log everything” on a transactional system), connection logging, shipped to a central SIEM with retention and integrity controls.
  • Evidence & target. Audit events visible in the SIEM; defined retention met; logs immutable off-host; a test change reconstructable to who/what/when.
  • Requirement. Configuration shall conform to a recognised hardening benchmark, verified periodically, with security patches applied within an SLA.
  • Our approach. Baselined against the CIS PostgreSQL Benchmark, remediated in risk order, with automated conformance re-scans and drift alerting; minor-version tracking with a restart-only update path and a reduced extension surface.
  • Evidence & target. Dated CIS conformance report with pass rate + remediation log; running version within the patch SLA (e.g. critical CVE ≤ 14 days); documented re-scan cadence.
  • Requirement. The database shall be network-isolated, reachable only from authorised hosts.
  • Our approach. A dedicated database network segment; firewall rules permit database and pooler ports only from the application and load-balancer subnets; the database listens on the internal interface only; no public exposure. The load-balancer health check probes the real database protocol, not a bare TCP port.
  • Evidence & target. A port scan from user networks shows the database port filtered; the ruleset permits only application-subnet sources.

3.9 Regulatory compliance (where a tender cites it)

Section titled “3.9 Regulatory compliance (where a tender cites it)”
  • Requirement. The solution shall support the tender’s compliance framework (e.g. GDPR, ISO 27001, national e-government security rules).
  • Our approach. The security controls above map to common frameworks: encryption in transit/at rest and access control (GDPR Art. 32; ISO 27001 A.8/A.10), audit logging and monitoring (A.12), and documented data handling. Right-to-erasure (GDPR Art. 17) vs immutable backups: immutable backups deliberately retain data (including deleted records) for the retention period; we document the legal basis and retention justification and reconcile erasure obligations with the retention policy up front, rather than claiming a conflict-free absolute.
  • Evidence & target. A control-to-clause mapping for the cited framework, plus the evidence artefacts (scan reports, drill reports, access reviews) that substantiate each control.

  • Requirement. The database shall sustain a stated concurrent load and meet a latency target without degradation.
  • Our approach. Connection pooling (PgBouncer, transaction mode) multiplexes thousands of client connections onto a small, bounded backend set, so the database is never overwhelmed by connection churn; memory and planner settings are tuned to the workload; latency is measured from actual query statistics.
  • Evidence & target. Sustains the stated peak concurrency with the pooler; committed p99 latency ≤ target from load tests with a representative workload; p95 stays stable as client count rises (pool wait ≈ 0 at peak).
  • Requirement. Checkpoint and WAL activity shall not cause periodic latency spikes.
  • Our approach. Checkpoints spread over the interval and made time-driven, WAL compression, and commit WAL placed on low-latency storage on a device separate from the data (and from the consensus store).
  • Evidence & target. No “checkpoints occurring too frequently” warnings; commit-latency p99 free of checkpoint-correlated spikes; throughput scales with concurrency up to the pool ceiling.
  • Requirement. Read-heavy workloads shall scale horizontally.
  • Our approach. A separate read endpoint fans read traffic across healthy replicas via the load balancer, offloading the primary. Consistency contract: replicas may serve slightly stale data (replication lag), so the application either tolerates lag on reads or pins consistency-critical reads to the primary — a documented application decision, not an implicit assumption.
  • Evidence & target. Read throughput increases with added replicas; the stale-read tolerance is documented per read path; replication freshness is monitored.

  • Requirement. The system shall preserve data integrity, prevent silent corruption, and prevent transaction-ID wraparound.
  • Our approach. Page-level data checksums, ACID transactions with constraints and foreign keys, full-page-write protection, autovacuum tuned above the conservative defaults with per-table overrides on hot tables, transaction-age monitoring, and periodic integrity verification of both live data and backups.
  • Evidence & target. Checksums enabled on 100 % of clusters; transaction age kept well below the wraparound ceiling (proactive alerting long before the limit); no emergency anti-wraparound vacuums; periodic verification runs with zero errors.

  • Requirement. The system shall be observable, with proactive alerting on health and SLO breaches, and the monitoring itself shall be monitored.
  • Our approach. Metrics from every database instance, the cluster manager, the consensus store and the operating system flow into Prometheus/Grafana; alerts route to a real on-call channel with a dead-man’s-switch so that a monitoring outage also pages; an independent synthetic database-protocol probe provides a liveness signal that does not depend on any single check. Alerts cover leaderless cluster, replication lag, connection saturation, transaction-age, disk and consensus-store capacity, backup failure, certificate expiry and clock skew.
  • Evidence & target. Mean time to detect ≤ 2 minutes for a fault; the alert set live and runbook-linked; the monitoring dead-man’s-switch continuously green. This directly addresses the failure mode where an unmonitored database can be down for days unnoticed.

  • Requirement. The solution shall be maintainable and patchable with minimal disruption, with reproducible configuration across environments.
  • Our approach. The cluster is defined as version-controlled Infrastructure-as-Code with an identical staging environment; minor patches roll out replicas-first via controlled switchover; major upgrades follow a rehearsed, backup-protected procedure; every production change is preceded by a staging run and followed by a verification record.
  • Evidence & target. Minor-patch write interruption ≤ ~10 s with zero read downtime (≥ 2 replicas), applied within a defined window of release; 100 % of configuration version-controlled; zero undocumented drift.

  • Requirement. The database shall have sufficient, monitored capacity headroom for the projected growth horizon.
  • Our approach. A capacity model driven by keeping the working set in memory, matching CPU to the pool-bounded active-query count, and provisioning storage IOPS/latency headroom on separated volumes (data, WAL, backups never share one thin volume); WAL and backup retention are treated as capacity inputs; the baseline is validated with load testing.
  • Evidence & target. Working set fits in cache (high cache-hit ratio); CPU/IOPS peak below a stated headroom threshold; storage alerting at 75/85 %; a capacity forecast covering the contract term. Prerequisite: the separated-volume and disk-sizing baseline must be met before this is claimed (on a constrained environment, call it out as a gating item, not an assumption).

  • Requirement. The platform shall avoid proprietary lock-in and interoperate via open standards.
  • Our approach. Open-source PostgreSQL with a standard SQL dialect and wire protocol, portable across on-premises, cloud and Kubernetes; migration via logical replication and standard dump/restore; no closed extensions on the critical path.
  • Evidence & target. The same deployment demonstrated on two substrates; export via standard tools; documented conformance to SQL-standard and open interfaces.

  • Requirement. The supplier shall provide defined support with severity-based response and resolution SLAs.
  • Our approach. Tiered support with a named escalation path and coverage hours (24×7 for the highest severity), a documented severity matrix, runbook-backed remediation proven against prior incidents, and post-incident reviews.
  • Evidence & target. A committed severity/response matrix (e.g. P1 response ≤ 15 min, workaround ≤ 4 h — set per contract), monthly SLA compliance reporting, and delivered incident post-mortems.

11. Load-Balancer Correctness (a differentiator worth stating explicitly)

Section titled “11. Load-Balancer Correctness (a differentiator worth stating explicitly)”
  • Requirement. The load balancer shall route clients only to a node genuinely able to serve its role, never to a dead backend that merely accepts a TCP connection.
  • Our approach. The load balancer uses application-layer health checks against the cluster manager’s REST API — a node answers “I am the primary / a healthy replica” only when it truly is. Bare TCP-port checks on database backends are explicitly prohibited, because a live pooler in front of a dead database would report healthy. A separate synthetic database-protocol probe provides an independent second signal.
  • Evidence & target. In a “backend down but port open” test the load balancer marks the server down and stops routing to it; the write endpoint always follows the current leader across failovers.
  • Why we call this out. We have seen a production database appear “healthy” for 12 days while completely down, purely because a proxy in front of it answered TCP and the health check was Layer-4. Getting this one detail right is the difference between a real SLA and a paper one.

Appendix — mapping to the ISO/IEC 25010 quality model

Section titled “Appendix — mapping to the ISO/IEC 25010 quality model”
25010 characteristicSections here
Reliability (availability, fault tolerance, recoverability)1, 2
Security (confidentiality, integrity, authenticity, accountability)3, 5
Performance efficiency (time-behaviour, capacity)4, 8
Maintainability (modifiability, testability)7
Portability (adaptability, replaceability)9
Functional suitability / compatibility (interoperability)9
(Operational) supportability & observability6, 10, 11

Companion engineering guide (the “how”, with commands): [[PostgreSQL-HA-Production-Install-Guide]]. Grounded in a real production PostgreSQL/Patroni cluster and the lessons in docs/CG-Stage/. Keep targets honest and evidence-backed — that is what wins technical evaluations. Last revised 2026-07-25.