CI/CD for Firmware When a Bad Build Reaches a Physical Machine
A precision contract manufacturer building embedded firmware for controllers, sensors, or production equipment is shipping something a typical CI/CD pipeline wasn't originally designed around: code that ends up running on a physical machine, where a bad build isn't just a bug, it's a potential safety and quality problem on the shop floor.
This checklist covers the pitfalls that matter most when GitHub Actions or GitLab CI is building firmware instead of a web application.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Firmware isn't software: what actually changes about testing it
A web application's test suite runs against the same environment it will deploy to. Firmware's test suite usually can't, because a standard CI runner doesn't have the physical hardware the firmware controls. Your pipeline needs a layer standard software testing doesn't: either a simulator that models the hardware's behavior closely enough to catch most bugs, or an actual hardware-in-the-loop rig the pipeline can drive directly.
Both GitHub Actions and GitLab CI can trigger a build and hand it to either kind of test layer; neither platform includes hardware simulation itself, so this is infrastructure you're building on top of the CI platform, not something the platform gives you out of the box.
Common pitfall: no hardware-in-the-loop test before shipping to a physical unit
A firmware build that compiles cleanly and passes a simulator can still fail on real hardware because of a timing issue, a sensor calibration difference, or a component tolerance the simulator doesn't model precisely. Route firmware builds through an actual hardware-in-the-loop test station as a required pipeline step before a build is approved for a physical unit, not as an optional check someone runs when they remember.
Set up a self-hosted runner physically near the test rig, since a hosted runner has no path to a piece of lab equipment sitting on your production floor.
How do you keep OT and IT networks separate in a CI/CD pipeline?
Manufacturing environments typically separate operational technology, the network running production equipment, from the IT network running everything else, for good security reasons. A CI/CD runner that bridges the two, by living on the OT network but pulling code from a cloud-hosted CI platform, can quietly undo that separation if it's not configured carefully.
Work with whoever owns your OT security policy before placing a self-hosted runner anywhere near production equipment, and keep the connection between the runner and the CI platform as narrow and outbound-only as the platform allows.
How do you tie a firmware version to a serial number?
When a physical unit ships with a specific firmware version, you need a durable record connecting that unit's serial number to the exact build, commit, and test results that firmware went through. Build this into your pipeline as an automated step, writing the build's identifying information to a manufacturing record system rather than relying on someone manually logging it after the fact.
This matters most when a firmware issue surfaces months after shipment: without that record, tracing which units are affected becomes a much slower and more uncertain process than it needs to be.
A pre-release checklist for firmware changes
Before a firmware build is approved to ship to a physical unit, confirm:
- Has the build passed hardware-in-the-loop testing on the actual test rig, not just a simulator?
- Is the build's version tied to a specific commit and test result set in your manufacturing record system?
- Does the runner that built and tested it maintain proper separation from your OT network?
- Is there a tested field-update or rollback path if an issue surfaces after units have already shipped?
A firmware pipeline's whole job is to prevent exactly the kind of failure that's expensive and slow to fix once it's out in physical units. Treat every one of these checks as non-negotiable rather than something to skip under a deadline.
Field updates: the rollback path most teams forget to test
Rolling back a web application means redeploying a previous version to a server you control. Rolling back firmware means pushing an update to units that may already be running in a customer's facility, on a network you don't control, sometimes with a person needing to be physically present. Test your field-update mechanism itself as part of your pipeline validation, not just the firmware build it delivers.
A firmware pipeline that produces perfect builds is only half the story if the mechanism for getting a fix onto already-deployed units is untested and unreliable when you actually need it. Include at least one full dry run of the update mechanism against a non-production unit in your regular release checklist.
What Good Looks Like
Good looks like every firmware build passing hardware-in-the-loop testing before shipping to a physical unit, with an automated record tying each unit's serial number to the exact build and test results behind it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
For the software and analytics side of a manufacturing operation, AWS gives you a deploy target for dashboards and quality data separate from anything touching the firmware itself.
For the IT side of the business, Vanta can monitor branch protection and access controls on your firmware repository, though it won't reach into OT-side hardware testing controls.
Frequently Asked Questions
Can we skip hardware-in-the-loop testing if our simulator is accurate enough?
Only for changes that don't touch anything the simulator has known limitations modeling, timing-sensitive code and sensor calibration are common gaps. For anything touching those areas, route through the real rig regardless of how good the simulator has been historically.
How do we keep a self-hosted runner near production equipment secure?
Place the runner together with your OT security owner, and keep its connection outbound-only where the platform supports it. Patch it with the same discipline you apply to any other machine on that network. Runner placement is a security decision, not a purely software one, so involve the people who own the production network.
What's the fastest way to trace which shipped units have a specific firmware issue?
An automated record tying each unit's serial number to its exact firmware build and test results at the time of shipment. Without that record already in place, tracing affected units after the fact usually means manually cross-referencing shipment logs, which is slow and error-prone.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
GitHub Actions vs GitLab CI vs CircleCI: Continuous Integration Comparison
Compare GitHub Actions, GitLab CI, and CircleCI: build speeds, runner pricing, matrix testing, Docker orchestration, secret management, and DORA metrics.
Cursor vs GitHub Copilot for Manufacturing Software Teams
Most of a manufacturer's code talks to a machine, not a browser. Where Cursor and Copilot fit MES and ERP integration work, and where neither belongs.
SOC 2 for Precision Contract Manufacturers
SOC 2 for precision contract manufacturers balancing shop-floor OT systems and office IT, and how Vanta, Drata and Secureframe fit each.
Application Security Where Software Meets the Shop Floor
Tradeoffs between Snyk and GitHub Advanced Security for a precision manufacturer whose software connects to ERP, MES, and machine-control systems.
Database Infrastructure for Precision Contract Manufacturers
Precision contract manufacturers integrating with plant-floor systems face different constraints than a typical software company. Here's the comparison.
CrowdStrike vs SentinelOne for Precision Manufacturers
Legacy CNC controllers cannot always run modern EDR. A step by step approach to the CrowdStrike vs SentinelOne decision for a precision manufacturing floor.