What the coding AI revolution isn't

There's lots of history for people to draw analogies from, but people have been doing so poorly.

Before I get going I need to lay some facts down. They're not fun facts, but they are facts. Lots of facts are objectively bad facts, but that doesn't change their fact-ness.

The nature of industrial software production has materially changed. This is a done thing. The coding LLMs have reached product market fit. It turns out that automating the language of automation was easier than you'd think. The cost of producing good enough coding LLMs is coming down sufficiently that open weight models are becoming viable (if less robust than the premium priced and rented frontier models) and pressure from Chinese model forges is further depressing token costs. Large scale software production is no longer viable using handcrafting techniques. Mid-scale software production is following suit. Industrial production of software is not the first industry to go through this transition, but it is the most recent and largest. This ship has sailed and our industry is in new waters whether we like it or not.

This factual statement says nothing about the ethics of the development of these tools. The ethics is where the large majority of the objection to these tools is founded, and there is some poor arguing going on about it.

But they don't even work!

No; but they are consistently bad in predictable ways, same as the transition from handcrafted machine parts to machined machine parts. The economics mean you get a lot more code for your dollar than handcrafting, and the rich guy who owns the code factory makes the call and not you. Same as it has been for the last 240 years of the industrial revolution. Software ain't special.

Now for the ethics.

The AI revolution isn't the sugar revolution

There are those who draw parallels to industrial sugar production starting in the late 1400s. This brought world wide change as calories were no longer scarce in Europe, the finance system of Europe stabilized, sugar stopped being an aristocrat-only food and started being available on street-corners, people started being able to have jobs that didn't involve agriculture thanks to the calorie surplus, which liberated enough mind and bodies from farmwork to lay the foundations for the scientific revolutions that led to alternate power sources, mechanization, automation, and ultimately computers. However, cheap sugar from 1470 to the early 20th century required chattel slavery.

A good book on this is Born in Blackness: Africa, Africans, and the Making of the Modern World, 1471 to the Second World War by Howard French (ISBN 978-1631495830, 2021).

The anti-AI argument drawing parallels to chattel slavery is doing so in a highly allegorical way. Here is this flagrant crime against human rights, see how it matches the crime of AI model training upon the property rights of IP owners. For one, this argument is patently offensive to peoples historically subjected to chattel slavery. For another, the property rights of IP owners is itself a modern invention post-dating the bust-up of chattel slavery in most of the world. This is an inappropriate argument, and should not be made.

Yes, these tools automate mining the commons for resale. The producers are even mining property without permission of their owners. Some got slapped for it (see the Anthropic settlement; my book was included in the haul but I wasn't in the settlement because my publisher didn't register the copyright) some are successfully wooing lawmakers to allow this little bypass of property rights for the good of everyone. And by everyone they mean everyone rich enough to own a hyperscale model forge. That should be challenged. Making parallels to chattel slavery is not how you do it.

The AI in everything boosters look at what sugar did to the world and see AI making that kind of change. Freedom from farming your own content! Expert analytical power at your fingertips! More time for fun pursuits, and less chasing obscure syntax errors! Boosters often overlook the cost of their futures. This is far from the first time humanity has seen a technological improvement, got all starry-eyed, kicked off a gold-rush, and created brand new toxic brownfields in the process to pollute generations (some late 1700s coal-ash heaps are still poisoning water tables).

The AI revolution isn't a new industrial revolution

Another common historical parallel is the industrial revolution built on the back of coal, and later on oil. The industrial revolution changed whole economies, providing work in cities for anyone tired of farmwork, increased the speed of travel which shrunk the apparent size of the world, increased interconnectednes of the whole world...

...and brought the current climate disaster we're living through. Training the big frontier models takes profound power and raw materials, exacerbating our current problems with decarbonizing our power-generation. Clearly, we don't need to make the same mistake twice?

This argument is based on our current US-centric technological model, which is being actively undermined from many quarters. The Chinese economy is somewhat more centrally planned, and is managing to both decarbonize and increase model forge capabilities in ways that simply aren't happening in the US. Europe is also making lots of noise about unwiring their technical dependence on US tech firms due to lack of trust in US privacy and security laws. The US tech companies are following US automakers in slowly weaning their global sales to focus on luxury sales domestically. This is a recipe for a major market realignment somewhere in here which will change the US economic calculus for trillion dollar hardware investments.

Second, we have known for two decades that Machine Learning models are much cheaper to train if you're training for a specific, targeted workload. We've had ML-based cancer detection for a while now. Building coding-AI is a targeted usage, and training open weight models are becoming viable as a result. Maybe not as deep-featured as the premier high-rent frontier models; but if you need that luxury experience, the US tech industry will happily take your money. Training for general purpose intelligence is far more expensive, but that market has not reached product market fit yet the way coding AI has.

Third, we have 60+ years experience watching industrial production reliant on cheap labor leave the US for cheaper shores. The US tech industry has been off-shoring industrial software production talent for decades whenever the company got big enough to economically operate in more than one country. Industrial production in other industries that has remained onshore tends to be one of: dependent on local distribution such as oil refining, fitting highly niche industries where talent isn't global yet, small-crafters who can't afford to offshore or are in it for the craft.

This is the same exact industrial revolution operating on a new industry, thatssit.

Eventually even the boosters will realize that limited model scope is the path to profitability, but it might take a stock market collapse to jog their memory.

The AI revolution is like any other industry transitioning away from hand-crafting

One historical parallel that has clear pedigree is the Luddite movement in the wake of increasing industrial automation changing the nature of work for whole industries of hand-crafters. This argument is closer to viable because it is founded in worker rights and not property rights. The rise of industrial automation in the 1780-1880 period also paralleled a rise in monopolies and crony capitalism where a few got fantastically rich while workers grew somewhat less skilled and less paid due to the reduced need for hand-crafting. It's also the era where Karl Marx did his writing. We're clearly in that spot again, and socialism is much less of a dirty word for US-based leftists. Reducing the need for highly trained hand-crafters in software will further reduce the aggregate quality of life for people involved with industrial software production, increasing the net-worth of the people owning the production pipeline with less of the RSU trickle-down we used to see with handcrafted software.

A parallel I haven't seen much is examining current model herding workloads. I bring you weather forecasting. In the age before computed forecasts, figuring out the weather was intuition, some basic math, and luck. In the modern era, official forecasters look at groups of different model outputs to build a consensus view, and then further tweak the consensus view due to intuition and experience. The last human review step provably improves forecast accuracy, which is why you want an expert there. Industrial software production is already experimenting with model-consensus views, but the cost of running multiple premier rented frontier-models in parallel is fast becoming too expensive for most.

This is its own damn thing

My last thought here: we still don't understand the talent economics of coding-AI. Right now we're riding the talent surplus of a bunch of expert hand-crafters retooling their skills for the model herding world, with a bunch more out of work right now acting as a latent pool to draw from. And yet, our educational pipelines are still tooled for producing hand-crafters. How do you effectively train consensus-modellers for code? Do we even know what a junior looks like yet? How do we train coding AI for new programming languages without significant existing codebases to train on? Does this mean new languages will take less guidance from human ergonomics and more on training ergonomics? So much is changing that when the latent talent supply dries up the entire industry will change once again to figure out how to hire for this. The labor market changes for industrial software production have only just started. 

But do not draw parallels to chattel slavery, even indirectly, or by allegory.

The people critique of AI

Nikhil Suresh of Hermit-Tech posted a critique of AI and how it affects decision-making today, and he brings up some points I want to highlight. Other people are tackling the technical and supply-chain sides of AI adoption in the tech industry; but this opinion is from someone acting as an Engagement Manager for a consulting firm, which puts them at the curious intersection of being both a Senior Product Manager and also a Sales Manager. What's more, Suresh is seeing this from the point of view of a team that can say no to AI strategy if they want (by not taking a contract). This means they see this issue as mostly a people and decision-making problem, which is a vantage point we haven't seen much from yet.

Those of us who have made it to Staff+ roles know that good software and systems are more the product of people than it is the technical foundations. Those technical foundations help improve velocity and maintainability, but bad people will make horrible systems out of excellent software. At the Staff+ level, a lot of our job is mediating differences between teams with different technical points of view, or informing managers of various types about the feasibility of different technical approaches. Sometimes we're listened to, sometimes we're not. As Suresh describes, we're in an industry-wide moment of sometimes we're not.

Suresh describes at length the biasing effects of C-level, whole company strategy to adopt a specific technical system in the absence of established best-practices. We've all heard the stories of TokenMaxing, teams yeeting AI-based workflows into deployment pipelines because it has to be done, and companies firing lots of people on the promise of AI only to hire a bunch back months later because the promise was not met. Some companies are doing the work; like Honeycomb, who has taken on the task of coming up with technically defensible ways of using AI tooling in the absence of a top down hype mandate. Most are not doing the work, they're providing massive incentives for adoption without an existing and proven framework.

This shows up in performance review and hiring cycles. News of annual performance reviews requiring evidence of AI use for "meets" or better were already being happening in 2023 when the tooling was much worse. By 2024 many companies required evidence you know how to use AI in coding workflows as part of their interview pipeline, something that has intensified in the last two years. Company strategy is hiring strategy, and company strategy was "we are now an AI company."

In a footnote, Suresh says:

If I was somehow CEO at a hospital or civil engineering firm, I would not for a second think it's my place to start mandating specific procedures or building techniques without explicit agreement from the professionals on staff - how fucking clueless are the non-technicians who have attended a few talks and are now making mandates about how their extremely expensive professionals are doing their jobs?

This is shedding light on a fundamental disconnect between tech CEOs and their workforce. Yes, we are highly trained; but that training was mostly paid by ourselves or other companies, and there are more of us being made every day. There isn't a Software Engineer shortage, and hasn't been one for over a decade. Labor costs are the number one cost center for making internet services, you bet your ass tech CEOs are going to invest in a tool that promises to radically reduce your main cost center. The big companies have been investing in off-shoring jobs outside of the US for over a decade specifically to reduce their production costs, AI came right for that desire and won.

Which is leaving us with an industry wide case of bad people turning excellent services into horrible services. Once the hype wanes (one company I know moved their big AI bet out of Manhattan Project mode and into a more sustainable Product Management cadence, there are others) we're going to be left with an industry wide hangover to clean up after. Neither will happen over night, but those of us who have been in the industry over 20 years (you don't have to be Staff+ for this) know what this looks like.

Growth vs throughput mindsets

I've talked about this one to various managers over the years when the topic of work-scheduling humans comes up. There are two big frameworks for deciding who works which tickets:

  • Throughput: Assign the person who will complete the task fastest to the ticket or story. Helpdesks often work on this model, since they're frequently quite under water.
  • Growth: Assign the second most knowledgeable person to the ticket/story (or third, or fourth if you have that many people) so they can get experience solving the problem. This method solves the cross-training problem.

Both approaches are valid, and which approach a team takes is balanced against the incentives the team operates in. For a Helpdesk team, which is often under-staffed, throughput tends to be highly prioritized. For a Dev team with on-call responsibilities, growth tends to be the norm to make sure problems can be handled by anyone and to stop burning out your senior talent. Security teams tend to be a blend of both; with ticket-based security reviews urging throughput mindsets, where incident-response prioritizes growth. For Platform teams that tend to be involved in a lot of incidents due to running the systems that problems happen on, even if the problem wasn't actually the platform system, throughput mindsets are hard to avoid.

Now add coding LLMs to the mix.

The on-label reason to use agents such as Claude is throughput: spend less time working on tasks to improve your throughput. For Engineering Directors this is a great thing, because you can get throughput advances in your reporting teams without sacrificing engineering maturity, letting you lean on the product roadmap a little harder.

I'd argue that the throughput nature of LLM agents actually sacrifices some growth at population-scale when used in currently typical tech-companies. By moving from a purely generative mindset when building code to a revisionary mindset, going from writing code to reviewing and editing generated code, engineers spend less time learning problem-spaces in detail. Less time learning means taking more calendar time building up true domain knowledge. Coding agents are a somewhat different thing than the decades of "just add another abstraction layer and forget about the lower level details" we've been doing since 1970. It is true that most engineers in SaaS companies aren't spending lots of time hand-tuning assembly for execution on their ISA, and doing so is actually very bad practice unless you're in extremely specific problem domains. The difference between yet another abstraction and coding agent is the difference between abstraction and synthesis.

Abstraction is a distillation of a problem-space into an API, which can represented by grpc protobufs or function-calls, or whatever. You have inputs, expected actions, and expected results; the abstraction handles the details so you don't have to and often someone else is responsible for updates over time.

Synthesis is building novel functionality through combining multiple things to create a new thing. Humans traditionally have been the synthesis engines of code-production, but coding agents are beginning to automate portions of this.

Writing code is synthesis. Until the advent of coding agents, your IDE would give you support through syntax highlighting and API short-cuts, reducing the cognitive load of bare syntax problems to let you focus on higher level problems like how a function should handle bad inputs. The IDE gives you support for abstractions, but leaves the synthesis to you.

Once coding agents get added in, you shift from generating code, synthesis, to reviewing automatically generated code whenever the agent creates something for you. This is when cognitive biases come into play. Remember the mindset thing I opened this post with? An engineer under throughput stress is going to be less diligent about checking generated code in detail, and will shift more of their domain learning to troubleshooting test-failures and incident-response. Less domain knowledge will be developed during the initial writing. An engineer with no throughput stress, perhaps they're doing some for-fun coding, is more likely to take a fine toothed comb to generated code to learn what the generated code is trying to do, why it works the way it does, and what best-practices it seems to be following; in this throughput-free case the coding agent is a driver of growth.

Coding agents end up magnifying the biases present in the team's environment, enhancing positive feedback loops. The advertised reason to adopt coding agents is to increase developer throughput! The throughput bias is baked into the technology and its marketing, which means organizations requiring coding agent use need to take steps to provide some negative feedback loop enhancement to avoid large-scale degradation in coding quality and incident severity. Use of these agents make an organization more sensitive to growth/throughput bias shifts prompted by org-chart and quarterly goal shifts.

Combine the bias-enhancing quality of coding agents with the industry-wide retraction in US-based jobs for economic reasons, and you should be seeing a general retraction in overall growth mindsets among US-based engineers.

The tech community needs so much repair

This BlueSky thread from Cat Hicks is on the money

There actually has not been enough reflection in the tech community about DOGE
all I see are "they're not engineers" "that's not what engineering is" I want a thick good really real piece to sink my teeth in about all the parts of this that ARE "like engineering"
I would write it except I want to stay safe
Engineering can be so much better than this and more than this. It's absurd how much we're resigned to this entire area of human work being held hostage by this kind of culture.
I do not believe that the majority of the software people I have worked with all these years sit around thinking, "I want cancer research to be destroyed." We have mountains to climb no doubt, we are trapped in some bad cultures no doubt, but I do not believe this of you for one moment.
Our tech community needs SO much more repair, processing, and collective reflection. The people who worked to support government and science were so betrayed by other groups. The people in flashy big tech cos feeling like their values were betrayed. Just so much repair needed here

I absolutely agree that our tech community needs quite a lot of collective reflection, processing, and work on repair. For a variety of reasons I've done a fair amount of casual research into recovering from trauma of various types. Some of this comes from having lived through the dot-com, great recession, and profitability-crunch contractions in the tech market, and some from good old fashioned life experience outside of the workplace. Trauma is trauma, and we have some pretty good ideas what chronic (ongoing, persistent) trauma does vs acute (single traumatic event) trauma.

Acute trauma: sudden-death layoff.

Chronic trauma: never being allowed to stay on one project long enough to get something good for your performance review, making you constantly fear the next layoff will have your name in the list.

Each of these affects the body and mind differently, and when entire populations are subjected to chronic traumas you get population-level reactions. When those chronic traumas are structural, as is the case with management culture in all of big US-tech, remediation becomes next to impossible. When individuals in this chronically traumatized population can't fix it, you get three big trauma-responses:

  • Cynicism. I can't fix it, that's just how it is. If you set your expectations low enough, you get to be happily surprised once in a while! It's great.
  • Heroism. Rally your fellow workers to overthrow the corrupt system! Who is with me? If not, I'll do what I can alone.
  • Trauma harder as a way of life. Obviously, we're in a cut-throat system which means I need to cut throats. QED.

Big-tech management likes workers in the trauma harder category, is somewhat tolerant of cynics, and is designed from the org-chart out to prevent the heroes from getting anywhere useful and to redirect their energies in positive directions like burnishing the company's reputation among diversity hires. Do this for a few decades and you have an entire population of highly educated workers who have been trained their entire careers to look at the next sprint/month/quarter's deliverable for your team and kinda ignore what the rest of the company is doing. How your quarter's project to reduce stream latency by 15% at the p95 level relates to the efficiency of data-analysis in ICE isn't always obvious, don't look up or you might find out.

So take these workers who live in this pit of oppression every day for their day-jobs, and put them into an open-source community for their fun-time activity. What happens then?

  • The cynics contribute as they're able, expecting corporate malfeasance to show up at any point, often seeing it when it isn't there.
  • The heroes go about building a community that actually is healthy for a change! Whew.
  • The trauma harder crew perpetuate corporate-style power structures because that's how tech works, accidentally reinforcing the cynics and frustrating the heroes

Fixing this sort of thing requires so much work.

  • The cynics need to be taught that their defensive pessimism is not appropriate by repairing the structural injustices
  • The heroes need to understand they're not alone and are being listened to through an effort of collective reflection
  • The trauma harder crew needs to realize that alternate structures are viable, and understand the damage they've endured in the existing system through extensive processing

You can't do this overnight, it will take a revolution of some kind. Some revolutions are slow, like the heroes getting somewhere with governmental support allowing unionization to creep in higher and higher numbers until union contracts dominate worker terms and conditions rather than Radford salary reports. Some revolutions come quick like whole industries getting nationalized after a socialist junta and remodeled away from oligarchic control. Some come generationally, like China overtaking the US for big-tech exports forcing the oligarchs to look elsewhere to stay fat.

Our tech community needs SO much more repair, processing, and collective reflection.

The desperation of Windows

It is no secret I'm a long time now-former Windows system administrator. The first Windows I professionally administered was Windows NT, and I was with it up through Windows 2008 (and a touch of 2012). I ran into an observation today that made me to Hmm, and that leads to blogposts.

@xgranade The common complaint is that you need to be a professional software developer to use Linux. But to use Windows 11 (and have it be usable) you need to be a professional sysadmin who uses terms like "group policy".

https://aus.social/@natarasee/115501020317612539

Because this is spot on. I'd argue Linux is usable by non-devs these days, but you still need a tolerance for fiddling and non-standard UI. The Windows side is extremely true. Windows in a corporate context is way more tolerable than Windows in a home context because the corporate context has a group of grumpy Windows sysadmins setting new Group Policy every time a security or feature release comes out to turn down the suck. Those grumpy sysadmins are as grumpy at Microsoft pulling this consent-violating shit as you are, and Windows lets you centrally shut it off (in a corporate context.) 

Windows has been losing desktop market-share to Apple for years, and the old "Wintel" cash-cow they used to enjoy is not milking as much as it used to. When software makers see flagging revenue and soft user demand, it's time to do demand forcing! And demand forcing leads to shittier experiences as the use more software! message gets ever more aggressive.

The long time followers of this blog have seen enough of this industry to know the cycle when they see it.

  • Darling product stops being darling for whatever reason. Competition, flagging significance, major incident spoiling user trust, private equity takeover, whatever.
  • The product's Product org has to make number go up in spite of all this so jacks renewal prices.
  • Renewal prices only jack so far before growth reverses, so Product has to ship new features to justify the price increases.
  • Bad uptake of new features means features get added to base plans to justify jacking the price.
  • Bad uptake of now baseline features means more aggressive prompting of those features to drive up Monthly Active Users metrics
  • Repeat

Do this for enough years and you get the sclerotic Windows 11, full of demand-forcing promptware that pisses your customers off.

On Fediverse, Paul Cantrell, CompSci professor at Macalester in St. Paul Minnesota, posted the following list:

Here’s the lightning sketch of Paul’s Treatise Against Efficiency that I’ve never written:

1. Efficiency is asymptotically inefficient: as costs approach zero, the cost of further reducing them approaches infinity.
2. Efficiency prioritizes the measurable over the difficult-to-measure.
3. Efficiency prioritizes what those in power see (or imagine) over on-the-ground reality.
4. Following from 2 and 3, efficiency reduces the amount and quality of information flowing into a human system.
5. Efficiency foments institutional inflexibility.
6. By removing slack, efficiency causes small failures to cascade more readily and increases the risk of catastrophic failure.
7. Following rom 4, 5, and 6, efficiency trades small costs for massive risks: from failures, from missed opportunities, and from inability to adjust.
8. Efficiency, when pushed, strangles the emergent phenomena that in the long term create all new things of value.
9. Thus, although it can be a by-product of evolution, efficiency as a goal in itself strangles evolution.
10. Efficiency as a goal strangles joy.

It turns out I do have time to write an article about that, below the fold.

ServerFault spam issues

For the maybe three of you still there, the StackExchange sites have been experiencing a multi-month spam wave from a certain group of spammers who've built automation for our sites. SuperUser gets it harder than ServerFault, but we both get hit. So I did some diving today to figure out just how bad the floods are.

For 2025:

  • Non-deleted questions: 2512
  • Deleted questions (mostly spam): 9748
  • One specific spam type: 6863

Yes, we're getting way more spam than legitimate questions right now, and have all year.

The Metasmoke people are doing a lot to make sure you rarely see it, but it's still a burden on the moderation staff due to the need to delete the spam-users. That deletion helps the StackExchange anti-spam systems identify bad IPs and others which increases the burden on spam-campaigners like these.

I wrote a weird little book. I'm still getting royalties, so thank you all for buying, but this book does not easily fit into 2025 concepts of "observability engineering," so I want to talk about my goals and how it still fits.

At the base, I ended up writing a book for Platform teams looking to deliver internally deployed observability systems. That's not quite what I had in mind when I started writing in 2020, but that's where it lives five years later. My actual goal was to write a book that was usable by people in the SaaS industry, but also in businesses where the main user of internally developed software was internal users. The non-SaaS population often gets ignored in book targeting, and I wanted something that would let these people feel seen in a way that reading yet another Observability for Cloud Systems book would. In 2025, this book is a Platform book.

In 2019 and early 2020 when I was working with Manning on the title and terms, the word "observability" came up. It seems hard to remember in 2025, but "observability' was still a vague term that didn't yet have industry consensus behind it. OpenTelemetry was a thing at the time, but the "metrics" leg of OTel was still in beta, and "logs" was merely roadmapped. In 2025 there are debates around whether the fourth pillar is profiles, performance traces, or errors, which could be stack-dumps or a category of logs. If we had decided to use "observability" instead of "telemetry" the book may have sold better, but the term "telemetry" works better for me because observability is a practice built on top of telemetry signals. I wasn't writing a book about practice, I was writing about herding signals.

Herding signals, not interpreting them.

In 2025, most of the herding is supposed to be done through OpenTelemetry these days. Or if it isn't OTel, the signals are being herded through other systems like Apache Spark. This is industry consensus; instrument your code, add attributes in the emitters and collectors, change your vendors as you need to, build dashboards in your vendor's platform. A rewrite of Software Telemetry would reference OTel far more often than I did, but I would still make sure to mention non-OTel styles due to OTel not actually being supported (or in some cases, a good fit) in certain environments like network telemetry.

Whatever the API format of the signals getting herded, platform engineers need to know the fundamentals of how telemetry systems operate and that's what I wrote about. But also, I wrote about storing those signals, which is something that OpenTelemetry deliberately leaves out as a detail for the implementer. As I extensively wrote about, storing signals and creating a reporting interface is a hard enough part of telemetry that you can build a business around it. In fact, the Observability Tools market in 2025 is valued at around $2.75 Billion, and they all would love for you to use OTel to send them data to store and present.

In the language of my book, OpenTelemetry is an early shipping stage technology. Early because it has no role in storage. OTel arguably has a role in the emitting stage through explicit markup in code itself. OpenTelemetry's impact to the presentation stage is mostly in tagging and attribute schemas and how they get represented in storage. Observability needs to consider every stage, but also the SRE Guide problems of figuring out what to instrument, to which markup standards, following which procedures to ensure reliability. Observability sits on top of telemetry.

One of the consistent comments I got during the pre-publication reviews was: "I want to know what to track."

My answer was simple: that's not the book I'm writing.

This book is for you, the growth engineer tasked with taking a Kafka topic (or group of topics) of logging data, sent there by OTel, and transform it in the big Databricks instance with all  the other business data.

This book is for you, the network engineer tasked with extracting network metrics out of a proprietary system, so you can chart network things in the main engineering dashboarding platform.

This book is for you, the security engineer tasked with extracting security event data out of a cloud provider to put into the SIEM system.

This book is for you, the project manager who has just been given a digital transformation project to revitalize how all the internally developed apps will produce telemetry, and how engineers will observe the system.

A sentiment just crossed Fediverse recently, which is in the vein of "RSS was peak social media, change my mind". The original post was from https://hachyderm.io/@Daojoan@mastodon.social and is quoted below:

RSS never tracked you.
Email never throttled you.
Blogs never begged for dopamine.
The old web wasn’t perfect.
But it was yours.

https://mastodon.social/@Daojoan/114587431688413845M

I was there for the rise and fall of blogging, so the rest of this post is me over thinking this particular post.

The Department of Government Efficiency, Musk's vehicle. made news by "discovering" the General Services Administration uses tapes, and plans to save $1M by switching to something else (disks, or cloud-based storage). Long time readers of this blog may remember I used to talk a lot about storage and tape backup. Guess it's time to get my antique Storage Nerd hat out of the closet (this is my first storage post since 2013) to explain why tape is still relevant in an era of 400Gb backbone networks and 30TB SMR disks.

The SaaS revolution has utterly transformed the office automation space. The job I had in 2005, in the early years of this blog, only exists in small pockets anymore. So many office systems have been SaaSified that the old problems I used to blog about around backups and storage tech are much less pressing in the modern era. Where we have stuff like that are places that have decades of old file data, starting in the mid to late 1980s, that is still being hauled around. Even when I was still doing this in the late 2000s the needle was shifting to large arrays of cheap disks replacing tape arrays.

Where you still see tape being used here are offices with policies for "off-site" or "offline" storage of key office data. A lot of that stuff is also done on disk these days, but some offices still kept their tape libraries. The InfoSec space is keen to point out you can't crypto-locker an offline tape, so offline tape is a useful tool in recovering from a ransomware incident. I suspect a lot of what DoGE found was in this category of offices retaining tape infrastructure. Is disk cheaper here? Marginally, the true savings will be much less than the $1M headline rate.

But there is another area where tape continues to be the economical option, and it's another area DoGE is going to run into: large scientific datasets.

To explain why, I want to use a contrasting example: A vacation picture you took on an iPhone in 2011, put into Dropbox, shared twice, and haven't looked at in 14 years. That file has followed you to new laptops and phones, unseen, unloved, but available. A lot goes into making sure it's available.

All the big object-stores like S3, and file-sync-and-share services (like Dropbox, Box, MS live, Google Drive, Proton Drive, etc) use a common architecture because this architecture has been proven to be reliable at avoiding visible data-loss:

  • Every uploaded file is split into 4KB blocks (the size is relevant to disk technology, which I'm not going into here)
  • Each block is written between 3 and 7 times to disk in a given datacenter or region, the exact replication factor changes based on service and internal realities
  • Each block is replicated to more than one geographic region as a disaster resilience move, generally at least 2, often 3 or more

The end result of the above is that the 1MB vacation picture is written to disk 6 to 14 different times. The nice thing about the above is you can lose an entire rack-row of a datacenter and not lose data; you might lose 2 of your 5 copies of a given block, but you have 3 left to rebuild, and your other region still has full copies.

But I mentioned this 1MB file has been kept online for 14 years. Assuming an average disk life-span of 5 years, each block has been migrated to new hardware 3 times in those years. Meaning each 4KB block of that file has been resident on between 24 and 42 hardrives; or more, if your provider replicates to more than 2 discrete geographic region. Those drives have been spinning and using power (and therefore requiring cooling) the entire time.

These systems need to go to all of this effort because they need to be sure that all files are available all the time, when you need it, where you need it, as fast as possible. If a person in that vacation photo retires, and you suddenly need that picture for the Retirement Montage at their going away party, you don't want to wait hours for it to come off tape. You want it now.

Contrast this to a scientific dataset. Once the data has stopped being used for Science! it can safely be archived until someone else needs to use it. This is the use-case behind AWS S3 Glacier: you pay a lot less for storing data, so long as you're willing to accept delays measurable in hours before you can access it. This is also the use-case where tape shines.

A lab gets done chewing on a dataset sized at 100TB, which is pretty chonky for 2011. They send it to cold storage. Their IT section dutifully copies the 100TB dataset onto LTO-5 drives at 1.5TB per tape, for a stack of 67 tapes, and removes the dataset from their disk-based storage arrays.

Time passes, as with the Dropbox-style data. LTO drives can read between 1 and 2 generations prior. Assuming the lab IT section keeps up on tape technology, it would be the advent of LTO-7 in 2015 that would prompt a great restore and rearchive effort of all LTO-5 and previous media. LTO-7 can do 6TB per tape, for a much smaller stack of 17 tapes.

LTO-8 changed this, with only a one version lookback. So when LTO-8 comes out in 2017 with a 9TB capacity, a read restore/rearchive effort runs again, changing our stack of tapes from 17 to 12. LTO-9 comes out in 2021 with 18TB per tape, and that stack reduces to 6 tapes to hold 100TB.

All in all, our cold dataset had to relocate to new media three times, same as the disk-based stuff. However, keeping stacks of tape in a climate controlled room is vastly cheaper than a room of powered, spinning disk. The actual reality is somewhat different, as the few data archive people I know mention they do great restore/archive runs about every 8 to 10 years, largely driven by changes in drive connectivity (SCSI, SATA, FibreChannel, Infiniband, SAS, etc), OS and software support, and corporate purchasing cycles. Keeping old drives around for as long as possible is fiscally smart, so the true recopy events for our example data is likely "1".

So another lab wants to use that dataset and puts in a request. A day later, the data is on a disk-array for usage. Done. Carrying costs for that data in the intervening 14 years are significantly lower than the always available model of S3 and Dropbox.

Tape: still quite useful in the right contexts.