nodramaops
Get in touch

Case studies

Eleven real stories from eight years. No internal company figures, just what changed for the business.

Case studies by track

Reliability under load

When every minute of downtime costs money and reputation.

High load2019 – 2021

45,000 requests a minute across several regions

A product with constant heavy traffic, where even a short outage hits revenue immediately. I led the company's move to containers and Kubernetes, rebuilt the release process and brought monitoring of every server in every region into one place.

Servers from every region on one screen, so problems are easier to spot early.

EdTech2024 – now

An online school with live video lessons

Kids and teachers meet in live video lessons every day. I'm responsible for the platform's infrastructure: video servers, databases, cloud, monitoring and incident response.

Changes can be checked in a separate test environment before release.

SaaS2021 – 2023

A design platform for users around the world

A geo-distributed platform running around the clock. I was on call for core services, rewrote the CI/CD pipelines, moved infrastructure to IaC and optimised cloud spend.

Infrastructure described in code, and fewer manual operations for developers.

Launches and migrations

When you need a foundation built, or a move made without breaking anything.

Product2023 – 2024

Infrastructure from a blank page

The company shipped mobile apps and a Telegram Mini App, but there was no infrastructure to speak of. I designed and built all of it: networks, servers, databases, automatic app builds and backend releases. Every service moved into containers with separate development and production environments, and users sign in through a single account system.

The company got infrastructure described in code, with monitoring, alerts and cost control.

Product2018 – 2019

A database with a standby in another zone

PostgreSQL ran in a cluster that the team had deployed and maintained themselves inside Kubernetes. I moved the databases to a managed service with a standby in another availability zone and set up backups of the whole cluster.

If one availability zone fails, the database switches to its standby, and the cluster can be restored from backup.

SaaS2021 – 2023

Builds that scale with the load

Builds and tests ran on a fixed set of servers. I rebuilt the process: build servers spin up to match the number of jobs and shut down when there is no work. Along the way I moved the WordPress site into containers and Kubernetes, and the packages from a self-hosted registry to a managed one.

Less waiting for builds at peak hours and fewer servers idling at night.

AI and automation

When the routine can go to a machine and people keep the decisions.

AITelegramOwn product

An AI bookkeeper for a 545-apartment building

Residents chip in for a backup generator and post about it in the group chat. Someone used to spend evenings reconciling it all in a spreadsheet. Now the bot reads the messages, screenshots and PDF receipts, matches bank payments against the statement, asks when something is missing, and answers "what was the money spent on in September?" with figures from its database.

Contributions for 545 apartments tracked with almost no manual entry, and bank payments matched against the statement.

AIEdTech2024 – now

AI assistants for an engineering team

I bring AI agents into engineers' daily work. Through MCP the agents see logs, metrics and the knowledge base, and help investigate problems.

Finding the cause of a problem gets easier: the agent gathers data from different systems itself.

Automation2021 – 2024

Bots that take routine off developers

A code repository bot, an automatic library update bot, and scripts that replaced manual operations.

Part of the routine people used to do by hand now runs automatically.

Security and control

When it matters who has access and who is knocking at your door.

EdTech2024 – now

A platform for people, not bots

I analyse suspicious traffic and build the defences: a web application firewall, rate limits and bot filtering.

More resources go to real users, because the firewall cuts off part of the suspicious traffic at the door.

High load, product2019 – 2024

Every permission under control

Moved public servers and databases into a private network, set up central access management and employee onboarding and offboarding.

Access is granted as needed and revoked during offboarding.