Reliability at scale (SRE)
Observability, SLOs, incident response, on-call, high availability. 30M requests and 100k jobs a day, a 500 GB database.
99.99%uptime, sole SRE
Platform Engineering & SRE
Vincent Legendre, Senior Platform Engineer / SRE. Eight years in the field, formerly a data and software engineer. Sole SRE of a platform serving 30M requests a day. I scope, I ship, I hand over — on my own.
Method
The right solution, not the most impressive one.
A controlled 30-minute cutover rather than a zero-downtime strategy, because the cost/benefit didn’t justify it.
Profiling, EXPLAIN plans, response-time distribution analysis, all the way to the patch.
Multi-provider architecture, with no dependency on a single supplier.
Infrastructure as Code, tests and observability from the outset, documentation as a matter of course. The team carries on without me.
Expertise
Eight areas, from reliability to the data platform — each backed by a result, not a claim.
Observability, SLOs, incident response, on-call, high availability. 30M requests and 100k jobs a day, a 500 GB database.
99.99%uptime, sole SRE
Front end, back end and database treated as one system. Failed background jobs down by 95%.
2×faster API responses
Application migration (Helm, reproducible deployments, controlled rollbacks, zero-interruption cutover), picking up a half-finished migration, running it in production.
3production clusters
Performance, high availability, replication, and the delicate operations done without downtime: version upgrades, pg_repack, Row Level Security with no regression.
0downtime on upgrades
Cloud-to-bare-metal migration, egress eliminated, at least matching performance.
3×lower infra bill
Full architectures: authentication, compute, databases, IAM, API exposure, WAF. From the network to application security.
Terraform, Pulumi, Ansible; secrets and configuration under control. CI bill cut to a third.
<3 minCI, from 10+
GCP-native data platform built end to end at a payments fintech (BigQuery, Dataflow / Apache Beam, self-hosted Elasticsearch, Airflow). Left the data team self-sufficient.
Tools
Orchestration & platform
IaC & CI/CD
Observability
Data & storage
Data engineering
Cloud & infrastructure
Languages
Engagement
A quick audit to size the work and put real numbers on the need.
Scope and pace set once the work is framed — never announced blind.
The deep understanding and development needed to ship. One change at a time.
I hand over: documentation and reproducibility. The team carries on without me.
Time and materials or fixed price, depending on how predictable the scope is. Fully remote.
Proof
Infrastructure bill cut to a third: compute moved to bare metal, egress eliminated.
At least matching performance.
p99 latency brought down tenfold — while the average never moved.
Uptime held for a full year, sole SRE, on a platform serving 30M requests a day.
100k jobs a day, a 500 GB database.
API response times halved; failed background jobs down by 95%.
Through instrumented diagnosis, not rewrites.
CI brought from over ten minutes to under three, CI bill cut to a third.
A scoping audit is enough to start: an objective diagnosis and a prioritised action plan, with no commitment beyond it.
Who it’s for
SaaS scale-ups, fintechs, high-traffic platforms, growing engineering teams with no established SRE function. Typically a team of around twenty developers that needs senior Platform / SRE expertise without hiring full-time.
An objective diagnosis and a prioritised action plan, with no commitment beyond it. Two fields are enough — the rest is quicker said out loud.
Thirty minutes, no commitment. Best if you already have a number that worries you.
See availability