DevOps consulting: what to ask before you hire someone

Rei Begaj· DevOps Engineer11 min read

A release takes most of Friday afternoon. One engineer knows how it works, and nobody is comfortable doing it while they're away. You could automate the deployment, but first someone needs to work out which steps are necessary and why the last attempt failed. That's a reasonable job for a DevOps consultant.

DevOps consulting covers the work around building, releasing and running software. Depending on the company, that might mean fixing a deployment pipeline or taking over a cloud platform. The difficulty when buying it is that very different jobs tend to arrive under the same heading.

Start with a recent release

Before asking for proposals, walk through a release with the people who did it. Where did they wait? Which steps required someone with special access? If the release failed, how did they get the previous version back?

You may find that the pipeline runs in ten minutes but approval takes three days. Or the tests pass, yet nobody trusts the test environment because it differs from production. Those problems need different kinds of help. Give a prospective consultant the details, including the awkward manual steps.

For measurements, DORA's delivery metrics are a useful starting point. They cover change lead time, deployment frequency, recovery after failed deployments, change failures and deployment rework. Apply them to one application at a time. A weekly release can be perfectly reasonable for one system while creating a backlog for another.

Agree on what improvement would be useful. For the Friday release, that might mean a second engineer can deploy during normal working hours and demonstrate a rollback. You can test that before accepting the work.

What belongs in the scope?

Ask the supplier to describe what you'll receive and what your own team will have to do. This matters especially where a proposal says something broad, such as “set up observability” or “implement infrastructure as code”.

Scroll the table sideways to see all columns.

WorkWhat to include in the proposalA useful acceptance check
CI/CDBuild and test steps, deployment permissions, secrets, approvals and rollbackSomeone on your team can deploy and recover a failed release
Infrastructure as codeVersioned configuration, reusable modules, state storage and an upgrade procedureYour team can recreate an environment and review a change before applying it
Cloud setup or migrationAccounts, networking, access, migration sequence and removal of the old setupThe application works after cutover, and the old resources stop accruing charges
MonitoringApplication metrics, logs, traces, alerts and retention settingsAn engineer can investigate a failed request using the information collected
ReliabilityAvailability requirements, backups, recovery procedures and incident responsibilitiesA restore has been tested and someone is responsible for responding to an alert
SecurityPipeline access, dependency checks, secrets handling and a process for fixing vulnerabilitiesThe team knows who can release code and who deals with a failed security check
Cloud costsCost allocation, unused resources, capacity choices and spending alertsYou can explain the bill by application or team and spot an unexpected increase

For monitoring, OpenTelemetry provides a way to collect and export telemetry across different tools. There's still work to do around it. Someone has to choose which events to keep, how long to retain them and which failures warrant waking a person up. Put those decisions in the scope too.

For security requirements, NIST's Secure Software Development Framework can help you prepare questions for suppliers. Ask how they handle a vulnerable dependency or a leaked credential in practice. A scanner's presence in the pipeline tells you little about what happens after it finds a problem.

If the proposal includes Kubernetes

Ask which part of your application needs it. Kubernetes can be useful when you need to manage many containerised workloads. It also introduces a cluster that someone must maintain, including its permissions, networking, updates and recovery procedures.

The CNCF's adidas case study gives a concrete example. adidas used Giant Swarm to help install and operate Kubernetes while its engineers concentrated on e-commerce development. The case study, published in 2019, reports that e-commerce releases went from one every four to six weeks to three or four a day. It also describes a platform team of 35 supporting about 300 engineers.

That staffing detail matters if you're considering a similar architecture. A company with a few applications and a small team should also price a managed container service or platform as a service. Ask the consultant to include the operating work in the comparison, particularly upgrades and support outside working hours.

The same question applies to an internal developer platform. It can let teams create environments and deploy without waiting for another team. But the templates need maintenance, and developers need help adopting them. DORA's platform engineering guidance recommends beginning with a minimum viable platform. Pick one task that developers repeatedly struggle with, make it easier, and see whether they use what you built before expanding it.

Who looks after each part?
Running the application
Cloud accounts
Access, networking and environments
Deployments
Build, tests and rollback
Infrastructure code
Configuration and state
Monitoring
Logs, alerts and retention
Security
Permissions and vulnerability fixes
Daily operation
Backups, updates and support

Agree which tasks the provider takes on and which stay with your team.

What will it cost?

Alpine Edge's DevOps engineering rate is CHF 100 per hour.

To estimate the work, we need to know which applications and environments are involved, what already exists and who can provide access. Ask for the expected hours by task and agree on a spending limit before implementation. Include reviews, documentation and handover in that discussion.

When comparing proposals, check who will actually do the work. Will the person designing the system also help implement it? How much time will your own engineers need to contribute?

Then look beyond the consulting fee. Build runners, managed databases and log storage have recurring costs. A migration may require you to pay for the old and new environments at the same time. Out-of-hours support may be a separate contract. Ask for a monthly operating estimate alongside the implementation price, with the assumptions written down.

A fixed price needs a defined scope: the applications and environments involved, the integrations, the migration sequence and the acceptance tests. If these are uncertain, a short assessment can give both parties enough information to quote the implementation. For work billed by time, agree how spending and remaining work will be reviewed.

Choose an arrangement your team can manage

For a one-off migration, a project with a clear end may be enough. If your team already owns the architecture but lacks experience with one part of it, an engineer working alongside them can be a better fit. Ongoing operations need a different agreement.

Scroll the table sideways to see all columns.

ArrangementUseful whenResolve before signing
AssessmentYou need to understand the problem before committing to a buildWhich systems will be examined, and will the findings include an actionable work plan?
Defined projectYou can describe the result and how to accept itWhat happens if a dependency or migration assumption turns out to be wrong?
Engineer working with your teamYour team needs specialist help or temporary capacityWho directs the work, and how will colleagues learn to maintain it?
Managed operationsYou want a provider to look after the platform over timeWho responds to incidents, what is covered outside office hours, and how can you leave?

Even with managed operations, keep someone internally responsible for the relationship. They need enough understanding to assess changes, approve costs and explain what the provider is doing.

If your own team will take over, start that work during the project. Give employees access to the repositories and cloud accounts, involve them in reviews, and have them carry out routine changes while the consultant is still available. Discovering at the end that only the consultant can deploy is an expensive handover problem.

Agree on how the work will be accepted

A timetable should leave room to discover problems. Migrating one application may reveal an undocumented database dependency; restoring a backup may expose a missing key. Ask where the plan allows for that investigation and who approves any extra work.

Demonstrate the design with a real application before moving the rest. Agree on the tests beforehand. Depending on the job, these might include deploying a change, rolling it back, restoring data and checking that access restrictions work. Use the results to decide whether you're ready for the next migration.

At handover, ask a member of your own team to follow the operating instructions. Let them find the relevant log, change a setting and explain the recovery procedure. The consultant can correct gaps while they're still on the project. Documentation is much easier to assess when somebody has to use it.

Sometimes the smaller job is enough

Moving an application to the cloud can be necessary because a data centre is closing or a hardware contract is ending. Be clear about that reason. If you also expect faster releases or less administration, identify the work that will produce those benefits. Copying the existing servers into cloud VMs may leave the same manual processes in place. DORA's flexible infrastructure research discusses migrations that leave the operating model unchanged.

Before committing to a large cost-optimisation programme, check whether you can account for the bill. Who owns each resource? Which test environments run all weekend? Which storage or logging policies were set once and never revisited? Those questions may give you a manageable first piece of work.

You can also stop after making deployments repeatable and recovery reliable. A small team may have little use for a service catalogue or a custom developer portal. Ask who would maintain it and how much time it would save the people using it.

Check the access you're giving the provider

A DevOps supplier may need access to source code, production logs, databases and backup systems. Find out who will use that access, where they work and whether subcontractors are involved. Agree how access is approved and removed when someone leaves the project.

For Swiss companies, hosting in Switzerland is only part of the data-handling question. The FDPIC's guidance on transfers abroad explains the conditions for transferring personal data to other countries. Check where logs, backups and support access go, as well as where the application runs. Requirements can also come from your sector or your customer contracts.

Where the GDPR applies, Articles 28 and 32 cover processor arrangements and appropriate security measures; Chapter V addresses international transfers. Have the relevant contract terms reviewed for your setup. EU financial-sector organisations may also have obligations under the Digital Operational Resilience Act, including oversight of ICT suppliers. This regulation shares the DORA abbreviation with the software research programme mentioned earlier.

If the platform will run AI systems, include those workloads in the review. Model access, retained prompts and deployment records can add responsibilities that a conventional web application doesn't have. Our self-hosted AI guide covers the hosting and operating decisions in more detail.

What should an assessment leave you with?

An assessment should give you enough information to decide what to commission next. Ask for a short map of the application, its dependencies and the route a change takes to production. Include the people in that map: a release that waits for one person's approval has a dependency too.

The findings should distinguish things the consultant checked from things they could not verify. “Backups are configured” and “we restored a backup successfully” describe different levels of confidence. Record missing access or unavailable test environments so they do not disappear from the implementation estimate.

A useful work plan names the first change, the reason for doing it, the expected effort and the person who will accept it. It should also say what can wait. If everything is urgent, ask which problem would still be worth fixing if you could fund only one piece of work.

Compare suppliers using the same example

Give each candidate the same recent release or incident to discuss. Ask them to walk through how they would investigate it, what access they would request and what they would show your team at the end. You will learn more from the questions they ask than from a long list of tools on a slide.

  • Ask to meet the engineer who would work with you, not just the person presenting the proposal.
  • Request an example of a handover document with client details removed.
  • Have the supplier explain one cheaper or simpler option and why they would accept or reject it.
  • Ask what happens when a dependency changes the estimate, or when you decide to stop after the pilot.

Write down the assumptions behind each quote before comparing totals. One may include migration rehearsals and training; another may leave both to your team. A short shared brief makes those differences easier to see.

Keep checking after the migration

Once the application is running, choose someone to review the bill, alerts and outstanding maintenance. A cost alert needs an owner who can explain the increase and decide what to do. A resource tag helps allocate a charge, but it will not shut down an abandoned test environment by itself.

Revisit the release example after a few normal changes. Can a colleague deploy without calling the consultant? Did recovery work when tested? Are people still copying secrets or changing production by hand? Use those observations alongside delivery metrics. If the new process is being bypassed, find out why before adding more automation.

Bring a specific problem to the first conversation

You don't need a complete technical brief to approach a consultant. A description of the last difficult release, a recurring incident or a cloud bill you can't explain gives the discussion somewhere to start. Include what your team has already tried and who would work with the supplier.

At Alpine Edge, we use a technical assessment to examine that problem and agree on the work before implementation. Our cloud and DevOps services cover the build and ongoing operation; for teams deploying models, we also work on AI integration and private AI. The first conversation should help establish what needs changing, what your team can handle and where outside help would be useful.

DevOpsConsultingCloud

Questions, answered

Alpine Edge charges CHF 100 per hour for DevOps engineering. The total depends on the agreed work and hours; cloud usage, licences and additional support arrangements need separate agreement.

A one-off migration or a gap in specialist knowledge can justify outside help. For continuing work, consider whether you need to hire internally. Either way, name someone on your team who can assess the work and make decisions about the platform.

That depends on your applications and the team available to run them. Ask what Kubernetes would solve in your case and have the supplier compare it with a managed container service. Include updates, recovery and on-call support in the comparison.

Potentially, subject to the applicable safeguards and any sector or contractual restrictions. Review where the supplier’s staff and subcontractors work, what they can access and where logs and backups are stored. A Swiss hosting address alone does not answer those questions.

A DevOps consultant helps a team build, release and operate software. The work can include deployment pipelines, cloud infrastructure, monitoring, recovery and access controls. Start with a specific problem and agree on a deliverable your team can test, such as deploying and rolling back one application.

Ask for an application and dependency map, a review of deployments and recovery, and a prioritised work plan with effort estimates. The findings should distinguish verified results from assumptions. A configured backup, for example, is not evidence that a restore has been tested.

There is no useful single duration for everything called DevOps. Repairing one pipeline is a different job from migrating several applications and their databases. After the assessment, ask for a phased schedule that includes access approvals, testing, migration windows and time for your team to learn the new setup.

It can be, especially when releases depend on one person or recovery has never been tested. Keep the first piece of work small: a repeatable deployment, working backups or clearer cloud costs. A small team also needs an operating setup it can maintain after the consultant leaves.

A provider can take on agreed operating tasks, but someone in your company still needs to approve costs, priorities and access. Specify incident response, service hours, escalation and exit arrangements. Buying a consulting project does not by itself provide ongoing monitoring or round-the-clock support.

Record a baseline for one application before changing it. Track delivery lead time, deployment frequency and failures alongside practical checks: can another engineer deploy, can the team restore data, and are alerts actionable? Review comparable periods and investigate the causes rather than treating a metric as a target on its own.

Keep repositories and cloud accounts under your control, with access to infrastructure state and keys. Require operating instructions and involve your team in changes during the project. Before handover, have a colleague perform a routine change and walk through recovery while the consultant is available.

Bring one recent release or recurring incident, a simple application diagram if you have one, and the relevant cloud bill. Explain what your team has tried and who can work with the consultant. An initial discussion does not require sharing passwords or handing over unrestricted production access.

Working on something similar?

Alpine Edge builds and runs this kind of system for clients across Europe and MENA. Tell us what you are trying to solve and we will tell you how we would approach it.

Talk to an engineer

Read next

All articles