Fleet automation
PowerShell running unattended across thousands of endpoints, wired into ticketing, orchestrated on a schedule. It runs at 2am whether I'm awake or not, so it had better be right.
That's how I explain my job when someone asks, and it's the honest version. I'm Zach — I build the automation underneath a managed IT fleet, and I love this work more than is strictly reasonable. Finding the thing everyone has quietly accepted as "just how it is" and deleting it is the best part of my week.
A few thousand machines, a few hundred businesses. Almost everything I've written exists because somebody was doing it by hand several hundred times a week, and that somebody was usually me.
The automation isn't really about the automation. Every hour a computer takes back is an hour somebody spends on something they'd actually choose.
Infrastructure engineering is a wide net. Here's where mine lands.
PowerShell running unattended across thousands of endpoints, wired into ticketing, orchestrated on a schedule. It runs at 2am whether I'm awake or not, so it had better be right.
Hyper-V builds, Server 2016 through 2025 lifecycle, P2V migrations, Active Directory upgrade paths, provisioning from bare metal.
Veeam design and fleet-wide upgrades. Nobody thinks about this until the one day it's the only thing in the building that matters.
HIPAA and CMMC environments where the evidence trail matters as much as the fix. Auditors have never once accepted "trust me, I did it."
I own the internal knowledge base. An automation nobody else can operate is a liability wearing a clever disguise.
Dashboards and tools people actually use to run all of it. Someone has to make the data legible, and I'd rather it be good. This site is a sample.
One script saves an afternoon. Thirty of them, running on a schedule across a fleet, stop being a convenience and start being capacity. Here's the shape of that, estimated conservatively.
Remediation, monitoring, deployment, audit, upgrade orchestration.
SOPs, runbooks, troubleshooting guides and escalation procedures — the knowledge base I own.
Work that happened without anybody starting it, checking it, or remembering it was due.
Most infrastructure people stop at the API. Most frontend people never see the fleet. I work on both ends, and the overlap is where the useful work is — the person who understands the data is usually the right person to decide how it should look.
This site is the frontend sample. Hand-written, no framework, one file, routed, themed, animated, accessible.
Three things I shipped, rebuilt so you can press the buttons. They run the real decision logic.
Anyone can write a script that works on their machine. The job is writing one that runs on three hundred machines you can't see, at an hour you aren't awake, without anyone needing to check on it. That changes how you build.
A hypothesis stays labeled as a hypothesis until something proves it. Root causes come from logs and command output, not from pattern-matching to the last thing that looked similar. When I'm wrong I say so out loud and name what changed my mind — I've retracted conclusions mid-project more than once, and the project was better for it.
I don't guess. If I say something is the cause, I can show you the proof. If I turn out to be wrong, I say so immediately rather than quietly changing my story — which costs me nothing and saves everyone else a lot of wasted time.
If one machine has it, I want to know how many do before I touch any of them. A fix applied to a symptom is a fix you get to apply again next quarter.
When something breaks in one place, I check how many other places have the same problem before fixing anything. Otherwise you fix the same thing over and over forever.
The happy path is the easy half. What it does when the disk is full, the service is already running, the credential expired, or somebody ran it twice at once — that's the half that decides whether anyone trusts it.
Making something work when everything is fine is easy. I spend most of my time on what happens when things go wrong, because that's the part that decides whether you can leave it alone.
Static analysis clean at warning level. Tests covering both paths. Documentation in the file itself. Signed commits. These run on my own personal repositories where nobody is checking, because a standard you only apply when watched isn't one.
Everything gets reviewed and tested before it goes live, automatically. I hold my personal projects to the same standard as my work, because a standard you only follow when someone's watching isn't really a standard.
I'd rather hold something an extra day than put a half-reviewed script in front of three hundred clients. Nothing in production has ever thanked me for being early, and a bad automation doesn't make one mistake — it makes the same mistake everywhere simultaneously.
I'd rather take an extra day than rush something onto hundreds of machines. A person making a mistake affects one customer. Automation making a mistake affects all of them at once, instantly.
Careful is the default, not the ceiling. When a security patch lands and every backup server in the estate is exposed, the right answer is a week, not a quarter — and the gating discipline is exactly what makes a one-week rollout safe enough to attempt.
Being careful doesn't mean being slow. When something genuinely can't wait, it doesn't — and the checks I build in are the reason I can move quickly without gambling.
I own the internal knowledge base, which means I've seen what happens when the only person who understands a system is on vacation. If I built it, somebody else can run it without calling me.
I write everything down. I've seen what happens when the only person who understands a system goes on vacation, and I refuse to be that person.
The work I'm proudest of didn't come from a ticket.
While working on something else I noticed a set of backup appliances that weren't being monitored at all. They held patient imaging. Nobody knew they were unattended because nothing was configured to say so. I raised it, then built the dashboard that watches them, with alerting that opens a ticket the moment one stops reporting.
A fleet-wide backup upgrade was scoped as written runbooks and a technician working through appliances one at a time. I wrote the automation instead. The manual guides were retired before anyone had to follow them.
An end-of-life migration across hundreds of endpoints and dozens of sites, with the scripting and the client coordination to make it land without a support queue full of people who couldn't work that morning.
Over a thousand internal documents — SOPs, runbooks, escalation procedures, troubleshooting guides. It started because I kept answering the same question, and became the thing the whole service desk runs on.
Careful is the default. It isn't the only speed I have.
Veeam released 13.1.1.18 to close a security vulnerability. Backup servers are the last line of defence for every client we protect, and doing them by hand would have taken weeks — weeks of exposure, with the most sensitive systems in the estate sitting unpatched.
I wrote the upgrade automation in a week and rolled it across the fleet. Every appliance patched, every machine that couldn't take it reported with the specific reason instead of being left in an unknown state.
Fast and careful aren't opposites. The preflight gating is what made a one-week rollout safe enough to actually do.
Open source unless a client's name is on it. All of it runs in production somewhere, which is the only review that counts.
Backup servers on a spread of versions, most not eligible for the new release. One convergent state machine instead of a script per scenario: preflight gating, config backup, restore-point baselining, checksum-verified staging, deferred reboot. Run it twice, get the same answer.
See the codeFinds dormant accounts still holding privileged access across two systems with no shared identifier. Distinguishes "no login on record" from "we haven't looked back far enough to know" — a difference that matters a great deal before you disable somebody's account.
See the codeHands the on-call phone from one engineer to the next. Writes every member explicitly, reads it back, fails loudly when the result doesn't match. Documents three API behaviors the vendor hasn't.
See the codeA site survey that used to mean a technician on site with a clipboard for most of a day. Now it's a zero-interaction collector that returns a structured report, so the engineering time goes into reading the findings instead of gathering them.
See the time it savesA golf club in an app. Handicaps to the real standard, leagues, live leaderboards, and a record of who played with whom. Built as a static site, grown into a real build with end-to-end tests, performance budgets and secret scanning, and shipped to phones.
Read the writeupThirty-odd pull requests to a public MSP script library — remediation, monitoring, deployment, audit. The long tail of a job where the same thing breaks at a hundred different sites.
See the pull requestsThree things I shipped, rebuilt so you can play with them. Each carries an efficiency grade — a weighted read on time returned, how hard it was to get right, and how many situations it covers.
Fleet upgrade — roughly 20 minutes of hands-on work per appliance, 265+ appliances, replaced by one scheduled run that reports what it refused to touch and why.
A network assessment used to mean a technician on site for most of a day, writing things down. Drag the slider to see what that costs across a book of business.
Network assessment — unstructured discovery turned into a consistent deliverable. Every site gets the same report, which is the part clients actually notice.
Our on-call rotation hands the phone over automatically. The phone system has an endpoint for exactly that, and a behavior that isn't in its documentation. Rotate the phone and watch.
Queue presence — saves minutes a week, not hours. It scores where it does because the failure it prevents is silent, and because the finding is now documented for everyone else who hits it.
I found this the way everyone finds it: two phones rang. The client that came out of it is open source, and I'm writing it up for the vendor's documentation.
Where things get broken before they go anywhere near a client.
Running today on a UniFi Dream Machine Pro with an 8-port switch, and a Synology NAS that hosts Proxmox and Jellyfin. That's the honest current state. The interesting version is in pieces on a desk, and I'd rather show you the finished thing than a parts list.
Boot straight to an IP and play Nintendo, PlayStation, or Xbox-era games off the network. Mostly an excuse to build something my friends will actually use.
DNS-level filtering for the whole house, integrated with the UniFi VLANs rather than bolted on beside them.
A Raspberry Pi picking up ADS-B and showing what's passing over the house. The answer is usually more than you'd think.
A Pi running the house. The ambition is a place where everything is observable, which is a strange thing to want in a home and completely obvious to anyone in this line of work.
VLAN layout, what runs where, and what I'd do differently. Written up here as it gets built, including the parts that go badly.
Waiting on a move before the full build-out. This page fills in as it happens.
Things I worked out the hard way, written down so the next person doesn't lose the evening I lost. Infrastructure, the homelab, the occasional round of golf.
An occasional note when I publish something. No schedule, no marketing, unsubscribe whenever you like.
I send these myself. Your address doesn't go anywhere else.
I'm Zach. I'm 27, I live in York, Pennsylvania, and I married Jordyn this year after ten years together. She's a dental hygienist, and I spend most of my days building for dental practices — so the software I'm fixing at work is the software she's using at hers. It keeps me honest about who's actually on the other end of a slow workstation.
I'm an infrastructure and integrations engineer. In practice that means I'm who gets called when something hard breaks, I own the internal knowledge base, and I spend most of my time taking work that someone does over and over and making it stop being work.
I didn't plan this. I started on a helpdesk answering phones. What I noticed was that I kept writing scripts instead of doing the same thing twice, and the scripts were better received than the manual work. So I kept going.
The thing I actually chase is pretty specific. It's the moment you find something everyone has quietly accepted — the weekly export, the manual check, the four-hour upgrade somebody does one server at a time — and you realize it doesn't have to exist anymore. Handing a team back their week is the best feeling this job has. I'd do that part for free.
I'm competitive, probably more than is healthy. Winning isn't really the bar. The bar is beating whatever I did last time. A shark has to keep swimming to stay alive and some days that's a little close to home, but it's also why the work gets better instead of just getting done.
I try hard to be honest about what I know. If I'm guessing, I say I'm guessing. If I was wrong, I say that too, as soon as I know. It's made me slower in meetings and a lot more useful in outages.
A tournament older than I am. I also wrote the app that kept score, which meant the math had to be beyond question before I could enjoy any of it.
My faith is the thing underneath the rest of it. The verse at the bottom of this site isn't decoration — it's what I come back to when a project is long and the finish line keeps moving.
Regularly, competitively, and with a putter I refuse to replace. It's the one place I can't automate my way out of being bad at something, which is probably good for me.
Same reason I like infrastructure — you can see how the pieces fit together. And nothing on a trail has ever opened a ticket at 4:55 on a Friday.
Constantly. Some of it technical, a lot of it not. Reading is how I keep finding out that a problem I'm wrestling with was solved by somebody else in 1974.
Always half torn apart. It's where I break things before a client ever sees them, and it's the reason I know what a bad idea feels like before I suggest it to anyone.
There's always something being built. If there isn't, I get restless and go find one. Jordyn can confirm.
Four years, four titles, and a web developer job that turned out to matter more than I expected.
Two and a half years split between systems work and building for the web. It's where the frontend half of what I do comes from, and the reason I don't treat an interface as somebody else's problem.
Information assurance and computer information systems, studied alongside working full time. Competed in the National Cyber League in 2023.
Teaching technology to adults who'd been told they were bad at it. The fastest way to find out whether you can actually explain something is to explain it to someone who has no reason to pretend they followed.
Hardware asset tracking and inventory across the client base, plus workstation provisioning for new onboardings. The ground-level view of how a fleet actually gets built.
First point of contact for dental and healthcare practices. A year and a half of learning what breaks, how often, and how people describe it when it does — which is still the most useful thing I know.
Supporting clients across dental, healthcare, government contractors and mid-market businesses throughout the Mid-Atlantic. Where the scripts started outnumbering the manual fixes.
Escalation point for hard infrastructure work, sales engineering scoping for software integrations and Azure infrastructure, fleet-wide automation, and ownership of the knowledge base.
Every one of those steps happened because I'd already been doing the next job for a while. I'm including the whole path rather than just the current title, because the helpdesk years are where most of what I actually know came from.
Contract, full time, or a question about something on this site. I answer everything.
Contract work, questions, corrections, or just to say hello. Whichever is easiest.
Emailzboogher@gmail.com LinkedInThe professional version GitHubEverything I've made public InstagramGolf, trails, and the lab