een a while, ey! I told you, blog writing is never going to be a regular thing. I keep it as a way to express my views and perspectives when I have consolidated all my thoughts and done my proper research to back them up. Anyways, getting to the point, as you remember, in that last blog I wrote about the analogy between AI and the industrial revolution. I also mentioned that it took a lot of time for the world to adjust to the effects of the revolution. Well, in the past couple of months, I realised the difference between that period and today's world is how fast the world has been adapting to change. For example, my LinkedIn feed spent most of its 2024 and almost 2025 telling me proudly that FAANG was dead. The same acronym termed by Jim Cramer on CNBC's "Mad Money" in 2013 that dominated the tech industry for several eras, that set of companies no longer hold the title of the big bull. Instead, the AI revolution has paved the way for its new big bull, MANGO (one of many acronyms that popped up and stuck with me): Meta-Anthropic-NVIDIA-Google-OpenAI. Well, the acronyms kept changing (sometimes adding an S at the end to MANGO to accomodate SpaceX) but they all radiated the same vibe, a major shift in power dynamics is right around the corner.
I started working on this blog a fortnight ago and kept putting it off. And in that one fortnight itself, Google signed a $920 million per month compute deal with SpaceX, SpaceX IPOed making Elon Musk the world's first trillionaire on paper, Anthropic launched Fable 5, their first publicly available Mythos-class model, just five days after publishing a paper asking the industry to slow down, the US government put a ban on Fable and Mythos three days after launch essentially shutting down their access to the whole world, and if all this wasn't enough, SpaceX, now being paid by both Google and Anthropic for compute, went ahead and cemented the deal to acquire Cursor, one of the most popular AI coding assistants among software engineers, for $60 billion. One fortnight of procastination, five MAJOR industry events. That's the pace we're dealing with.

Its now mid June 2026, and I've been tracking the AI industry for the past several months. While I do not feel any drastic shift in the industrial tectonic plates, I do feel some rumbling in the industry. So I thought to share with you, the Good, the Bad, and the What Now. This blog is going to be a long one to grab a coffee and lets get right into it.
The Good
After years of the industry essentially brute-forcing the rising AI demand problem since the launch of ChatGPT in 2022, by throwing increasingly expensive hardware at it, the complacency period bill has finally caught up. The software requirements have caught up to the hardware innovation done so far, and the industry is now forced to actually innovate rather than just scale.
One of the most fundamental limitations of running large AI models (majorly talking about LLMs) has always been memory. Traditionally, a model loads its entire weight into memory before inference. This meant running capable models required expensive hardware capabilities that the task might not even need. This was improved with the introduction of MoE (Mixture of Experts) architecture wherein different tasks were delegated to models of different capabilities and sizes based on the task complexity but it was still a workaround, not a solution. But the open source community, specifically through llama.cpp, solved this. They did this by enabling layer-by-layer loading wherein only the parts of the model actively needed were in the memory at any given time. For those complexity fanatics out there, it was like refactoring a polynomial nested loop into separate linear loops that are now processed sequentially giving out the same output with the fraction of cost. This architectural shift opened the door to running even more capable and demanding models on consumer hardware.
Talking about one of another open source community win (love how humanity still has the capability to collaborate so efficiently even when times feel bad, love you guys!), the open source community has been quietly closing the gap on proprietary models, giving them a run for their money (which I feel they do not have, more on that later). One of the most significant moment was DeepSeek R1 in Jan 2025. Built in two months, under $6mil, using lower capacity chips, it matched the performance that big tech had spent billions achieving. Since then, DeepSeek V4 delivers performance rivalling top closed-source models at roughly ~27% of the compute cost of its predecessor. There have also been several such open source models like Kimi K2.6 by Moonshot AI which tie GPT 5.5 on SWE-Bench Pro at roughly 80% lower costs per token. That "80-90% of the results with 10% of the cost" documented argument is soon going to be essential while considering the cost of running this tech (wait, am I seeing another analogy here: similar to the industrial revolution having the evironment costs as side effect, the AI revolution having the tech costs as its side effect)
Talking about costs, the cost reduction and peformance optimisation means nothing without the tooling to actually run these models locally. Luckily there's a true saying, "If you're thinking about it, someone potentially already thought of it". And the open source community never dissapoints. The local model hosting ecosystem has also been maturing rapidly with the introduction of Ollama for those TUI/command line nerds, LM Studio for those non-developer normal people, Jan for those privacy fanatics, vLLM for those production-grade maniacs and lastly GPT4All for those who want to run models on a CPU without a GPU at all. The last one is totally not my favourite just because I can't afford anything above GTX 1650. But in all seriousness, it shows how the open source community has their own not-so-little world independent and equally, if not more, capable than the big tech.
Talking about big tech, a few weeks ago, Google Research published TurboQuant, a vector quantization algorithm that compressed the runtime memory of LLM models during inference. I am not going to force bore you with the details, you can check them out in the reference below but the improvement numbers are significant, and that too without any model retraining or weight modification, that's a huge deal. Making AI cheaper to run is the kind of "no hurrays" work the industry currently needs. And publishing a research paper on it feels exactly like a full nostalgia moment created by Google Research considering it was Google Research itself whose 2017 paper on transformer architecture "Attention is all you need" kicked off this whole LLM and AI craze starting with ChatGPT in 2022.
Talking about one more win on the proprietory side, NVIDIA recently launched RTX Spark, not just a new chipset which "might" be able to run GTA 6 on 60 FPS (about 60 years later when GTA 6 finally lauches on PC), RTX Spark is NVIDIA's first full SoC (System on Chip), built in collaboration with MediaTek and Microsoft. 20 CPU cores, 6144 Blackwell architecture GPU cores, and 128GB of unified memory (over-compensation much). NVIDIA's direct and spiritual answer to Apple Silicon, it's a tightly integrated architecture that brings serious compute in a compact, power-efficient form factor. Now time will tell whether the benchmarks hold up in the real world.

One last one. This one particularly hits close to home not only because I've been following this guy for more than a decade, but also because this story tells something deeper than what seems on the surface. On May 31 2026, PewDiePie, the famous YouTuber with over 100 million subscribers, open sourced Odysseus, a fully self hosted AI workspace that runs entirely on your machine. It has a built-in tool called Cookbook that guides you and lets you choose which models runs best on your hardware. It includes OS-level agents capable of managing files, mailboxes, calendars, local tasks, to name a few. But thats just its functionality. This represented something which not many ventures did till now. This project by a YouTuber, who held the most subscribed channel on YouTube for years, was the first YouTuber to reach 50 million subscribers, whose 2019 beef with mega corporation T-Series forced YouTube to change its subscriber counter from a live accurate display to a rounded, abbreviated display, that hit 30,000 GitHub stars in 48 hours. While I haven't tested Odysseus myself (I don't even have the excuse of not having the required hardware this time), this showed one key detail. All prior open source tools reached developers. This reached EVERYONE. The framing he used at launch, "war on big tech has just begun" doesn't feel like a baseless claim or a hypothetical anymore.
The Bad
Now, even if the good section makes you feel optimistic, I wouldn't cash in on those thoughts if I were you just yet. Every coin has two sides. Let's turn our coin upside down.
Amazon had introduced an internal AI usage leaderboard called KiroRank in May 2026 built on top of the company's Kiro AI developer platform to encourage employees to adopt AI tools. Great. But what actually happened was it only measured sheer consumption, employees gamed the system and started doing absolutely anything to climb the rankings. This "tokenmaxxing" (love the internet) got out of hand pretty quickly and literally by the end of May 2026, upper management had to step in, sunset KiroRank and also lost a significant AI budget in the process. They did however shifted to tracking "normalised deployments" or the AI-generated code that actually ships.
Speaking of budgets, these couple of months were full of such corporations publicly being reported to be facing similar issues where AI at such a scale costs more than anyone has ever thought about. Microsoft for example, introduced Claude Code internally in December 2025 but had to backtrack six months later in June 2026 when token-based billing consumed their annual AI budget well ahead of schedule, and they had to cancel most of their licenses. Uber also deployed Claude Code to roughly 5000 of its engineers and burned through its entire 2026 AI budget of a whopping $3.4 billion (yup, thats a huge capital B) in just four months. Thats roughly $500 to $2000 monthly API costs per engineer. The shocking thing is that these examples aren't startups that are miscalculating runway, or who forgot to add limits to their AI spend (oh and btw this is real and it wasn't a small team, it was an unnamed Fortune 500 company who spent a whopping $500 million on Claude AI in one month), or small non-technical team memebers like a Product Manager who doesn't know how to efficient use AI and accidentally spend a whopping ~$1400 worth of compute in just one hour due to an AI agent going rogue and getting stuck in a loop of "reloading context" or "self-correction (this happened too). These are two of the most operationally sophisticated companies on the planet, with over a combined workforce of 262k professionals and NOBODY saw this coming.
And its not just these miscalculations creeping within the costs, its the infrastructure as well which is not able to keep up with the demand, IN 2026. If you're a technical person who has been been using AI coding assistants like Cursor or Claude Code, let me know if you've felt this. I know people using Cursor and I myself have been using Google's Antigravity, and its the same. For me, as of May 2026, every third prompt has been a very productive "servers are busy, please try again" error. For Cursor, its "Our servers are currently overloaded.." error. But the idea is the same. Well this can be understood for a 2022 startup called AnySphere's AI IDE Cursor but for a 1998 company like Google with involvement in literal cloud services via its Google Cloud Platform (GCP), that's very astounding. And it wasn't just me, Reddit threads are full of the same complains across all users, lowest non-paying users to highest plan paying users. And not just Antigravity and Cursor, almost every major model provider has been just playing catch-up trying to cover the costs of these plethora of services they have been providing, stretching themselves thinner and thinner with every second and every user while simultaneously failing to serve the demand of what they've already committed to.
And it's not that these corporations are just watching their ship sink or just stopping these services altogether. Naah! They would never. Instead what is taking a toll is the flat subscription model. These AI services are slowly shifting from a flat cost for a AI service subscription to a usage based pattern (and spoiler alert, that too isn't working). Google's Antigravity previous offered three model quotas, a smaller allocation for premium models like Claude (Sonnet and Opus) and GPT (GPTOSS), a mid tier allocation for their own Gemini Pro, and an almost unlimited allocation for Gemini Flash. Clever users (like myself, I know right!) learned to plan and architect with premium tiers and implement with the Flash. But since the launch of Gemini 3.5 Flash at Google I/O, the three categories collapsed into two, the smaller allocation of premium Claude and GPT models and the mid-tier allocation of Gemini models (Flash and Pro squished in one but with a fraction of what it used to be at the same or higher price). Google did triple the paid plan quotas twice in response to user backlash and also reset everyone's limits but it still wasn't close to the original. GitHub Copilot simultaneously shifted from the subscription to usage-based billing. The message from every provider is essentially the same, we cannot keep selling unlimited access because unlimited access at this scale doesn't have a sustainable economical model in the long run.
You would argue, look at the industry as a bigger picture. Look how much money is flowing. Looking how NVIDIA employees became millionaires ever since the AI boom between 2023 and 2025. Look how SpaceX and its employees also had the same fortune of becoming millionaires when it IPOed a few days ago. How every major AI company is easily raising 100-200 rounds of funding every second, and getting evaluated at a bajjillion dollars worth even before IPOing (don't know what will happen when companies like Anthropic, after raising so much money, will be evaluated at when they IPO then in Fall of 2026). Money is there man, it's just taking time to catch up. Well, can't believe I'm about to say this, you're right ...... and also wrong! There is money, but it's the same money circulating through the market giving the illusion of a bajjillion dollars. Step back from individual company raising money stories and look at how capital is flowing. NVIDIA sells chips. AI companies buy those chips and also receive investment from the same chip makers. Cloud providers buy more chips and also invest in AI labs. AI labs spend their funding on cloud compute from those same providers. The same dollars are going around the same track and are on their 2498274th lap. There's a very interesting bloomberg blog that perfectly sums up EVERYTHING I'm trying to say (even where I got to know about this)
At the end of the day, you can see the cracks everywhere. OpenAI is projected to lose around $14 billion in 2026, nearly triple its 2025 losses, even as it projects $100 billion in revenue by 2029 (these "losess" and "revenue" words play a crucial role in these sentences). According to the SEC filing, Alphabet, a company that spent 15 years buying back its own stock, just issued equity for the first time since two decades to raise $80 billion for AI infrastructure, pushing its total debt past $100 billion. Anthropic confidentially filed for an IPO at a near trillion valuation. And now, just recently, Google is paying SpaceX $920 million PER MONTH for compute in a three year deal. While researching about this one, I got to know, Anthropic, a company heavily backed by Google, is also paying SpaceX $1.25 billion PER MONTH for the same reason. The bubble keeps getting bigger and you know the end result of such a bubble.
(EDIT) Jun 12: Well, the more I try to shorten my blog, the more content I get every day, I swear this will be the last edit. On June 9th, Anthropic launched Fable 5, their first public available Mythos-class model, just five days after publishing a paper asking the industry to slow down on AI development. Well, three days after launch, June 12th, the US Department of Commerce issued a legally binding export control directive ordering Anthropic to suspend all access to Fable 5 and its restricted sibling Mythos 5 for any foreign national, INCLUDING Anthropic's own foreign national employees, REGARDLESS of where they are located. Considering that Anthropic has no way to track users by nationality in real time across dozens of cloud platforms simultaneously (talking about stretching thin like a rubber band), it shut down both models for everyone. WITHIN 3 DAYS. This was the first government-forced takedown of an AI model in history. And by the looks and pace of it, it doesn't seem to be the last. There's a whole ordeal of why this happened and how several reports indicate that it was the researchers of Amazon, one of the major investors behind Anthropic, who were the reason behind it but sadly, it has now started to feel like a normal Tuesday. One observer on Fortune put it best, "If you describe your product as a munition in every press release, eventually a government takes you at your word. They wrote the legal predicate themselves and called it a brand."
What Now?
Told you this was going to be a long one. I bet you never sat with a coffee to read it, right? Aah, that's fine! I'm sure this blog was atleast 3-4 coffees long. Or very close to what many would say as someone who "rambled on for 18 pages. FRONT and BACK!". Well, I know how this "Good" and "Bad" coin must have you confused at the current state of the industry. Don't worry. That's literally everyone, even the investors of this expanding bubble. You might argue that it just might be a string of unfortunate events. Might be true, but only if you look at each one individually. Zoom out and its a shape or pattern you've already seen before (another analogy warning!)

Think about it, LinkedIn feed from the past 2 years telling you FAANG is dead, telling how the new order has arrived, complete with a fresh new face and acronym. MANGO. MANGOS. Whatever. A major shift in power and you need to reposition yourself. Well, I did a bit of digging (so you don't have too, you're welcome) and you could see it in the builder/lower layers too, not just corporates. Hackathon submission queues, YC batch compositions, startup focuses all shifted visibly year over year as the craze took hold. I would say hackathons and incubators are a more clearer indicator than a now social media platform like LinkedIn because they reflect what people genuinely believe and bet on, and not just for virality. And when both layers align, you know the wave is real. But guess what, this has happened before.
Web/Dot-com: ~1995-2000: Everything needed a website. No website, you're a loser. App Developement: ~2008-2015: The launch of iPhone in 2007, the App Store in 2008 with 500 apps. Within years, every brand, franchise, and startup needed its own app. Every online instructor and every course started to teach app development. "There's an app for that" became a badge for cultural relevance. No app, you're a loser. Cloud Computing: ~2010-2018: Everything and everyone needed to be online. AWS, Azure, GCP became the new infrastructure religion you need to follow. You hosting everything locally, what a loser! (ok, I'll stop) For me, this was an understandable craze given how the connected everything was aiming for but most absurd endpoint for me for the craze for cloud gaming. Google lauched Stadia in 2019 with enormous fanfare and shut it down in Jan 2023, refunding every purchase. NVIDIA GeForce Now survived but still. IoT: ~2014-2019: Amazon launched the first Echo in November 2014 and started the smart speaker craze. By 2018, almost half of US consumers owned a voice-activated speaker. Every product, lights, plugs, fridges, thermostats, toothbrushes, fridges, everything needed to be "smart". The energy peaked around 2019 and quietly plateaued. Most of those still exist and sell but the craze is gone. ML/AI/Deep Learning/Neural Networks: ~2016-2020: Geoffrey Hinton's deep learning breakthrough had been sitting since 2012 but it was around 2016 when the gold rush truly hit mainstream. Suddenly every product needed an AI feature, every startup pitch deck needed "powered by machine learning" somewhere on the slide, every hackathon project had a classifier that nobody asked for. TensorFlow dropped in 2015, scikit-learn became every data science course's best friend, and if your app didn't have some form of "intelligent" feature, good luck at demo day. The irony of course is that this era is what directly laid the tracks for the LLM craze that followed. Crypto/NFT/Web3: ~2019-2022: The NFT market tripled in value in 2020 to $250 million. In 2021, the average NFT sale price hit $2044, the highest ON record. Then major cryptocurrencies like Bitcoin fell more than 70%+ from their highs and by mid 2023 over 95% of NFT projects had effectively zero trading activity. Web3 and crypto still operate behind the scenes for decentralisation or security but the hype was ridiculous. AR/Metaverse: ~2021-2024: In October 2021, Zuckerberg renamed Facebook's parent company to Meta and declared the company would become "metaverse-first". Reality Labs, the division in Meta responsible for it, recorded a cumulative operating losses of approximately $83.6 billion from 2020 through 2025. Apple dropped Vision Pro in 2024 for a whopping $3500 and almost nobody bought it. LLMs/AI: ~2022-present: I mean ..... "What! You wanted more?!?"

The pattern is not lowkey. The industry finds a new tech, plays it on loop until exhausted, extracts maximum value, and inflicts maximum collateral damage, then jumps ship to the next track when someone puts on something new. Sounds kinda like my Spotify journey. Every cycle follows the same shape: genuine innovation => industry-wide pile-on => unsustainable economics => painful correction => the tech finally finds its real footing, quieter and more durable than the craze suggested. Here's something worth marinating in your head. The foundational transformer technology behind every LLM you've used was developed by Google's research team published in a 2017 paper called "Attention Is All You Need". Google just didn't believe in it enough to act on it. OpenAI did. The next technology that redefines everything might already exist in a paper that nobody is paying enough attention to (Or maybe someone is in some research lab).
There's one more thing coming that is being talked about but not loud enough yet. Every AI model is judged on two basic metrics: inference speed and output quality. Until now, the output quality hasn't been the bottleneck, the internet since the last three to four decades provided vast amounts of human generated training data and models improved rapidly. The industry's focus has therefore been almost entirely on inference speed, either through hardware improvements like RTX Spark or software optimisations like TurboQuant or llama.cpp. But here's what's slowly happening. The internet is now filling up with AI-generated content. Code, images, videos, writing, comments, everything. Which means the next generation of models will increasingly be trained on the output of previous models. You've probably seen the experiment people did with AI image generators: take an output, feed it back in as input, repeat. The image degrades with every generation. Errors amplify. Fine details dissolve. It looks fine in the beginning phases, you had to do it many times before the degradation becomes visible. Or think about a photocopier. You take a photocopy of a photocopy of a photocopy and so on and watch the quality degrade. But nobody does that with a photocopier on purpose, you go back to the original. The difference with AI training data is that we don't have that control. The original is the internet and the internet itself is being overwritten in real-time. We are polluting our own original piece. While the corporations are trying to filter for human-generated data, the competitive pressure to ship new models fast rather than precisely makes careful data curation a luxury. I thought this was my original thought but it even has a name - model collapse and it's being studied and early signs are already documented. When the data quality problem becomes undeniable, the industry will be forced into its next phase, not faster inferences, not cheaper compute, but usably quality data ingestion and curation. And in my opinion, that's when the technology matures. A boring, but a valid correction that follows this craze.
There's a quote that I came across recently that I think cuts deep
You'll never know what the next song is and when it will come to light. But the pattern says someone, somewhere, is already writing it. And if history is any evidence of the little analogies it puts, it's already sitting in a research paper that an organisation doesn't quite believe in yet. But don't you worry,
