In 2026, David Holz described a fairly abnormal scene for a product launch. He went to Google and asked for 10,000 GPUs for Midjourney's first week. Not for some hypothetical scale five years later. For week one.1

A few minutes later in the same conversation, he gave the other half of the story: Midjourney had no investors. Holz had decided to start with his own money and said he kept spending it until he was roughly $1,000 in the negative.1

Those two facts sound almost incompatible. On one side, infrastructure at the scale of a major company. On the other, a founder who says he emptied his own account rather than raise a venture round.

That contradiction is a useful way into Midjourney's origin story.

Today, the service is mostly described through outcomes: millions of users, an enormous Discord community, images with an instantly recognizable aesthetic and revenue estimates measured in hundreds of millions of dollars. A February 2026 article on Andrew.ooo goes as far as assigning Midjourney $500 million in 2025 annual revenue with 163 employees, or a little over $3 million per employee.14

The number is spectacular. It is also less solid than the legend built around it. Midjourney is private and does not publish detailed financial statements. Several figures about the company get repeated from site to site until they start to look like audited accounts. The same Andrew.ooo article, for example, says Holz's previous company Leap Motion had raised more than $300 million. TechCrunch counted roughly $94 million by the time it was acquired.12 The article also presents Midjourney as publicly available by February 2022, while Discord's own account still describes a private beta followed by the public opening in July.10

There is a more interesting story here than that of a magically efficient startup.

It starts before Stable Diffusion. It passes through researchers releasing models, artists experimenting in public notebooks, and a Discord interface that first looks like a shortcut and then turns into an enormous machine for collecting feedback. And it contains a gap Midjourney has never filled publicly: nobody outside the company knows exactly what its earliest models were or what their training cost.

Before generative images, Holz was already trying to make interfaces disappear

David Holz did not arrive at Midjourney as a computer-vision researcher who had spent ten years training image models. His previous obsession was almost the inverse problem.

At 17, for ISEF 2006, he built an array of microphones able to locate a voice in three-dimensional space. He later studied applied mathematics, worked at NASA Langley and the Max Planck Institute, then left his PhD program to co-found Leap Motion.1 15

Leap Motion tried to replace part of the physical interface to computers with hand tracking. Its small black sensor watched your fingers and translated their movement into input. For roughly a decade, Holz worked on a question that is easy to state: how do you make the interaction between a person and a computer more direct?

The commercial result was much less elegant than the demos. Leap Motion raised nearly $94 million according to TechCrunch, reached a $306 million valuation at one point, struggled to find a sustainable market and was acquired by Ultrahaptics in 2019 for roughly $30 million.12

It is tempting to turn this into a neat anti-VC fable: first company takes venture money and disappoints, second company bootstraps and wins. Reality is less tidy.

We do not know Holz's cumulative compensation at Leap Motion, how much of the company he still owned at the sale, the liquidation preferences attached to investor shares, or the details of any Ultrahaptics retention package. We therefore do not know how much personal capital he actually had when he began Midjourney. His 2026 account establishes that he used his own money and eventually exhausted it. It does not establish the starting amount.1

What he explains much more precisely is why he changed problems.

After ten or twelve years working on hand tracking, Holz says he realized that a better interface was not enough. Even perfect hand tracking would not necessarily help someone think, find an idea or understand what they wanted to make. He wanted to move one level up and investigate the creative act itself: where ideas come from, how we explore them and how we work with other people to turn them into something.1

Midjourney therefore did not begin as “a startup that wanted to generate pretty pictures.” Holz describes a small exploratory organization with several projects around imagination and creation. Images won because a new family of models arrived at the same time.

2020-2021: diffusion becomes a usable toolbox

The word “diffusion” feels ordinary now. In 2020, it was not the obvious engine for a mass-market creative product.

The best-known generative image systems of the period were still strongly associated with GANs, networks where a generator and discriminator improve against each other. GANs could already create striking images, but training could be temperamental and controlling them with free-form language was a separate problem.

In June 2020, Jonathan Ho, Ajay Jain and Pieter Abbeel published Denoising Diffusion Probabilistic Models. The simplified idea is strange but elegant: train a network to reverse a process that gradually destroys an image with noise. At generation time, begin with noise and denoise it step by step until an image appears.18

In May 2021, Prafulla Dhariwal and Alex Nichol showed at OpenAI that diffusion models could beat leading GANs on several image-synthesis metrics. Crucially, code and checkpoints followed. OpenAI published guided-diffusion, including pretrained models up to 512 × 512 pixels under an MIT license.6 7

Between those two moments came another decisive component: CLIP.

OpenAI released CLIP in January 2021. The model learns to bring images and natural-language descriptions into a shared representation. It is not an image generator by itself. It is closer to a system that can measure how well an image corresponds to an idea expressed in words.5

That was enough to ignite a remarkable period of creative experimentation.

If a generator produces an image and CLIP can evaluate how closely that image matches a phrase, the score can become a compass. Artists and developers began combining VQGAN, CLIP, diffusion models and different guidance methods in public notebooks. The images were slow, unstable, often bizarre and sometimes beautiful. For the first time, a free-form sentence could feel like a practical way to navigate a visual space.

Holz himself names one person in an August 2022 interview with The Register: Katherine Crowson, known online as RiversHaveWings. He says her work on CLIP-guided diffusion was the development that particularly caught his eye.2

That matters because it prevents two opposite mistakes in the Midjourney story.

The first would be to present Midjourney as an isolated invention emerging from a sealed proprietary lab. It did not. Holz explicitly connects the company to breakthroughs around transformers, CLIP and diffusion, and to community experiments that made those ideas tangible.2 3

The second mistake would be to conclude that Midjourney V1 was simply “Katherine Crowson's notebook with branding.” Nothing public establishes that.

The more accurate answer is less satisfying: we know the intellectual family of early Midjourney, not its recipe.

What we actually know about V1

In the 2026 ISEF conversation, Holz revisits the moment when he focused on diffusion. He says he was reading the papers, encountered diffusion models and felt they were fundamentally different from GANs and older neural-network approaches. He discussed them with researchers he knew around San Francisco. In his retelling, the reaction was essentially: this is a big deal, somebody should work on it.1

He chose to work on an unusually poorly defined area: creative images.

That fits his background. A metric can tell you reasonably well whether a model recognizes a cat. It is much worse at deciding whether an image is interesting, beautiful, surprising or useful for unlocking an idea. Holz says he is drawn to these “fuzzy” problems, the ones without a single clean number to optimize.1

Midjourney's current documentation gives a surprisingly precise chronology of its early models:

VersionPeriod as default modelInitial grid resolutionMidjourney's current description
V1February to April 2022256 × 256very abstract and painterly, low coherency
V2April to July 2022256 × 256creative, colorful and painterly, low coherency
V3July to November 2022256 × 256highly creative compositions, moderate coherency
V4from November 2022512 × 512new codebase and new architecture designed by Midjourney

Those dates come directly from Midjourney.4

They also kill one persistent explanation: the first Midjourney versions cannot simply be fine-tunes of the publicly released Stable Diffusion 1.x weights.

Stable Diffusion's public weights were released on August 22, 2022.8 By then Midjourney had already gone through V1, V2 and part of V3. There can still be common research influences, exchanges among researchers and similar components. The Latent Diffusion Models paper also existed before Stable Diffusion's public release. But the timeline rules out the common story in which Holz downloaded public Stable Diffusion and added a Midjourney aesthetic on top.

The timeline still does not tell us what was actually running behind /imagine in February.

Midjourney has not released V1 source code, a V1 checkpoint, an architecture diagram or a training log. When The Register asked Holz about the technical stack in 2022, he declined to detail it. He would only say that the company had large AI models with billions of parameters trained over billions of images.2

That is roughly the documentary boundary.

We know CLIP-guided diffusion directly inspired Holz. We know strong public diffusion code and pretrained weights were available in 2021. We know Midjourney was training its own large models in 2022. We do not know whether V1 began from a specific OpenAI checkpoint, a model trained internally from scratch, a multi-stage mixture of components, or an architecture already heavily rewritten by the Midjourney team.

Going further would turn a lineage into imaginary leaked source code.

Open research lowered the price of experimentation, not the cost to zero

That distinction also matters for the most tempting question: how much did the first Midjourney training run cost?

The public answer is frustratingly simple: we do not know.

I could not find a primary source assigning a training cost to V1, V2 or V3. There is no published GPU-hour count, cloud invoice, run duration or exact training-cluster size.

The 10,000-GPU figure Holz gave in 2026 does not solve this. In his account, he asks Google for those GPUs for the week-one launch. He is explaining the difficulty of making diffusion models available to large numbers of people, not reporting the accelerators used to train V1.1

Mix those two things together and you can create a wonderfully dramatic cost estimate in about thirty seconds. It will also have very little evidentiary value.

We can instead use another model from the period as a benchmark, as long as the word “benchmark” stays attached.

The later PixArt-α paper compares its own training cost with Stable Diffusion 1.5. Its authors estimate SD 1.5 at 6,250 A100 GPU-days, roughly 150,000 A100-hours, for an estimated $320,000 of training compute.13

That is not the cost of Midjourney V1.

Stable Diffusion 1.5 is not Midjourney V1. Architecture, data, resolution, training schedule, cloud provider and possible reuse of pretrained components differ. A successful final training run also does not equal total R&D cost. Failed experiments, data preparation, storage, engineers and all the inference used to test the product still exist.

But the comparison gives one useful scale. By 2022, serious text-to-image training could already cost hundreds of thousands of dollars in compute without automatically requiring the tens or hundreds of millions now associated with frontier language models.

Early Midjourney could have cost less than a full Stable Diffusion training cycle if it reused pretrained components. That is technically plausible in such an open ecosystem. It remains an inference until Midjourney publishes the actual training recipe.

So the honest answer is less exciting than an invented spreadsheet: the exact cost of Midjourney's first training runs is not public, and the available evidence is not sufficient to reconstruct it reliably.

The training data is less mysterious

On another part of training, Holz was much more direct.

When Forbes asked in September 2022 how Midjourney's dataset had been built, he described a large scrape of the internet and the use of publicly released open datasets.3 He also confirmed that individual consent had not been obtained from living artists or copyright holders whose work appeared in those sources.3

In another interview, when asked directly about the ethical considerations of training these models on publicly available images, Holz does not defend the dataset through a theory of consent or compensation. He moves the question toward how the tool is used.19 He brings up an AI system recreating Doom and recounts that John Carmack, the game's co-creator, did not object to a model learning from his work. Holz then uses a deliberately simple analogy: if somebody invents a hammer, we do not normally object when other people use it to drive nails.19

The answer is revealing precisely because it does not really answer the accusation that AI companies are stealing images. It shows Holz's underlying position instead: he treats AI as a new general-purpose tool and evaluates it mainly through its ability to solve problems and expand what humans can do.19 That leaves the concrete questions behind the controversy unresolved: who had the right to place images in a training corpus, where learning ends and reproduction begins, and whether creators should be able to opt out or receive compensation.

In other words, Midjourney has been relatively explicit about the broad origin of its data and much less persuasive when explaining why that origin alone would make the training legitimate. Holz does not deny learning from works available online; he implicitly disputes the premise that such learning is itself an illegitimate appropriation.3 19

That exchange matters for two reasons.

First, it shows how many different things can be hidden inside the word “open.” CLIP code can be public. A diffusion checkpoint can be downloadable. A dataset can have its own license. A commercial model can then be trained on an assembly of open resources and web-scraped images. All of this is frequently compressed into “based on open source,” as though one license described the entire stack.

Second, it shows why Midjourney could not rely on raw training images as its only long-term advantage. Any sufficiently funded competitor could also collect huge numbers of images. Midjourney needed to learn something that other companies did not possess in the same form.

That is what V3 begins to provide.

Discord was not just a lazy interface

From a distance, launching an image generator inside Discord can look like the choice of a team too small to build a proper web product.

There was surely an efficiency benefit. A Discord integration lets you avoid immediately rebuilding accounts, chat, sharing, community, notifications and a social layer. But Midjourney's user testing also found that its initial AI-enthusiast audience preferred Discord, according to Discord's own case study.10

The choice changed the product itself.

In a conventional image app, you type a prompt, wait, inspect the result and repeat. In Midjourney's public channels, you continuously see everyone else's generations, their words, mistakes, variations and selections. Use becomes observable.

Discord describes a small private beta, followed by an invitation system where paying users could invite friends. In July 2022, Midjourney opened the server to the wider public. Within three months of private testing, the app hit Discord's previous one-million-user server limit.10

The growth eventually forced Discord to change its own infrastructure. In an engineering post titled “Maxjourney,” Discord later described the challenge of supporting Midjourney after its community passed ten million members with more than one million people connected at the same time.16

The platform choice gave Midjourney several things simultaneously: an interface, organic distribution, a permanent live demo and a community that learned to use the product in public.

Most importantly for a model lab, it produced preferences.

V3: when users begin training the advantage

When a user receives four images, they are not just consuming a result. They make a selection.

They upscale one. Ask for variations of another. Ignore the remaining two. Rewrite the prompt. Try again. At sufficient scale, those actions create an extraordinarily rich record of what people find appealing or useful.

In August 2022, Holz told The Register that V3 was the first version to incorporate a feedback loop based on user activity and response. He emphasized that the striking improvement had not come from simply adding more art to the dataset. It came from data about which images users liked and how they used the system.2

That is the point where Midjourney's story changes.

At the start, the lab benefits from a wave of open science: CLIP, diffusion papers, public code, checkpoints and community notebooks. A small team can move quickly because it does not need to reinvent every foundational component.

Then the product begins to generate a resource that does not exist elsewhere in exactly the same form: the preferences of Midjourney's own users.

Holz returned to the idea in 2026. He described Midjourney as a large model taking in dozens of signals from millions of people, and talked about feedback loops intended to amplify collective exploration rather than merely converge on the same results.1

That is probably a better description of Midjourney's moat than “its pictures look nicer.”

The aesthetic is visible. The loop that keeps producing it is not.

V4 is the clearest documented technical break

Midjourney released V4 in November 2022. For the first time, its own documentation uses unusually explicit language: an entirely new codebase and brand-new AI architecture designed by Midjourney, trained on a new Midjourney AI supercluster.4

A few months later, Google disclosed the infrastructure detail that had been missing. Midjourney trained the fourth version of its algorithm on Cloud TPU v4 using JAX, while running inference on GPUs.9

That still does not reveal the architecture. It does show a transition.

V1 through V3 emerged in a period when Midjourney spoke openly about community research inspirations while already protecting its exact stack. V4 was presented as a new internal architecture trained on very large-scale infrastructure. Between those moments, the company had acquired users, revenue, preference data and a reason to depend less on the components that had made its first experiments possible.

Open source was not the final product. It was the launch accelerator.

That trajectory is common in technology. Midjourney just moved through it unusually fast: less than a year between a 256-pixel V1 from an exploratory period and a V4 the company already described as an architecture of its own.4

Financially, Midjourney was almost the inverse of Leap Motion

While the technology became more proprietary, the financial structure remained remarkably simple.

In August 2022, Holz told The Register that Midjourney had around ten people, was self-funded, had no investors and was already profitable.2 His description of the business model was almost brutally direct: image generation is expensive, users pay the cost of using it, Midjourney adds a margin, and that margin has to feed the team and fund more work.2

That is not the standard generative-AI playbook, where a company first raises a very large amount of capital to buy compute and only later works out what customers are willing to pay.

Leap Motion helps explain the preference. In 2023, The Information reported that Holz had soured on venture capital after his first company raised heavily and ultimately sold for far less than its peak valuation.11 Midjourney rejected investor approaches and instead put a large share of its revenue back into the expensive chips required to train and run its models.11

Holz's 2026 account pushes the story further. He says he was not fundamentally anti-investor. He simply wanted to use his own money first. Then he used it all.1

That constraint does something interesting: research has to meet a paying product very quickly.

A well-funded lab can spend years treating revenue as a future problem. Midjourney did not have that luxury. When inference is expensive and the founder's own capital is heading toward zero, “do people enjoy this enough to pay?” stops being a presentation metric. It becomes the lab's power switch.

There is an important caveat to the bootstrapping myth. An unknown founder with $1,000 of debt does not normally ask Google for 10,000 GPUs and get a call back. Holz arrived with a decade of industry history, technical credibility, relationships and the ability to contact researchers and infrastructure providers.1

Midjourney's bootstrap did not happen in a vacuum. It relied on social capital, prior experience, industry access and an open research ecosystem. That makes it more useful to study than the inspirational-poster version.

Revenue: directionally clear, numerically fuzzy

Midjourney's financial growth is impressive enough without decoration.

In 2023, The Information reported that the company was on pace to exceed $200 million in revenue with about 40 employees, no outside capital and a profitable operation.11 In 2026, Forbes' company profile cited PitchBook for $300 million in total revenue in fiscal 2024, again describing Midjourney as profitable and still without outside funding.17

After that, the figures become softer.

Andrew.ooo gives $500 million for 2025 and 163 employees, which produces the headline ratio of roughly $3.07 million per employee.14 The estimate appears on many sites, but I could not find a primary financial document or independent investigation at the level of The Information that locks both values down.

It is more useful to treat the figures as different confidence levels:

MetricPublished figureConfidence for this article
2023 revenue> $200M / reported annual pacehigh, The Information reporting
2024 revenue$300Mreasonably high, Forbes citing PitchBook
2025 revenue$500Msecondary estimate, use cautiously
2023 headcount~40high, The Information
2025 headcount163secondary estimate, sources diverge
VC raised by Midjourney$0strongly documented by Holz and multiple sources

The important point survives the uncertainty. A company training its own image models reached hundreds of millions of dollars in revenue with an unusually small team and without a conventional venture-capital round.11 17

That is already strange enough. It does not require an extra zero added by folklore.

“Zero marketing” needs an asterisk too

Another phrase follows Midjourney around: the company supposedly built all of this with “zero marketing.”14

As a literal accounting number, I could not find a primary source proving that marketing expenditure remained exactly zero across multiple years. Midjourney has a brand, community work, partnerships, events and a public presence. Which of those costs land in an accounting line called “marketing” depends on how the company books them.

The more interesting idea behind the claim is well documented: distribution was built into the product.

Discord says Midjourney grew organically from its private beta, first through invitations and then through the public server. The app eventually escaped its own community: Discord says the Midjourney app has been installed in roughly seven million servers.10

Every public generation shows the product to somebody else. Each image is simultaneously a result, example, accidental tutorial and sometimes an advertisement. Users learn prompting by watching other users. Images spread around the wider web without Midjourney needing to buy the space where they appear.

That is not “no distribution.”

It is distribution where part of the acquisition cost is absorbed by normal usage.

And that same usage sends better data back into the model. The economic loop and the training loop begin to merge.

The Midjourney loop

We can now reconstruct the origin without filling its gaps with mythology.

In 2021, several advances become usable at the same time. CLIP connects language and images. Diffusion models jump in quality. Researchers publish code and checkpoints. Artists such as Katherine Crowson show that these components can become creative instruments driven by text.2 5 6

Holz arrives with a question rather than an architecture: how do you expand human imagination? He recognizes diffusion as a mechanism suited to a domain where “good” and “bad” cannot easily be reduced to one objective function.1

A small team builds early models. Their exact recipe remains secret, but their scientific lineage is visible. V1 is already the default model in February 2022. V2 follows in April. The product grows inside Discord and opens its beta widely in July.4 10

The service then has to solve a problem the research notebooks did not: serve huge numbers of generations in parallel. Holz says he asked Google for 10,000 GPUs for the week-one launch at scale.1

Users start paying quickly enough to cover compute and leave a margin. That removes the immediate need for a funding round. More customers finance more infrastructure. More usage creates more preference signals. Those signals improve V3. A better model attracts more users. Revenue funds the supercluster and V4, which Midjourney explicitly describes as a new internal architecture.2 4 9

The loop looks like this:

open research
      ↓
diffusion + text prototypes
      ↓
social product on Discord
      ↓
paying users
      ↓
compute funded by the product
      ↓
user preferences and behavior
      ↓
better tuning / new models
      ↓
more usage
      ↺

That is where Midjourney becomes hard to copy.

A paper can be reproduced. A checkpoint can be downloaded. An image-generation interface is not mysterious. Rebuilding years of user choices, an enormous creative community, a continuously tuned aesthetic and infrastructure financed by recurring customer revenue is a different problem.

So what model is Midjourney actually “based on”?

The answer changes depending on the period.

For the earliest experiments, Midjourney clearly belongs to the 2021 CLIP + diffusion ecosystem. Holz directly cites Katherine Crowson's CLIP-guided diffusion. OpenAI's public code and checkpoints show the sort of technical raw material available to a small team at the time.2 6

For V1 through V3, the exact architecture is not public. Midjourney was training its own large models and already refused to detail the stack.2 It is therefore misleading to call them simply “open-source models,” just as it is misleading to present them as public Stable Diffusion fine-tunes.

For V4, Midjourney explicitly says it built a new codebase and new architecture trained on its own supercluster.4 Google confirms TPU v4 + JAX for training and GPUs for inference.9

For later versions, the system becomes still more proprietary. Midjourney discusses capabilities and goals publicly, but says far less about weights and architectures.

The short answer is therefore: Midjourney was born from a wave of very open research and code, then quickly evolved into proprietary technology about which the company discloses relatively little.

That is not a contradiction. It is probably the central pattern in the story.

And how much did the first training run cost?

After all this research, the answer is still: we do not have the number.

We can at least separate what is known from what is not:

QuestionWhat the sources support
Was V1 based on public Stable Diffusion?not in the ordinary sense: V1 predates public Stable Diffusion weights by months
Did CLIP influence Midjourney?yes, directly documented
Did CLIP-guided diffusion influence Holz?yes, he explicitly names Katherine Crowson
Was Midjourney training its own models in 2022?yes, Holz describes large internal models but not the architecture
How many images were used for V1?unknown
How many GPUs trained V1?unknown
What did V1 training cost?unknown
Were the “10,000 GPUs” the training cluster?not documented that way: Holz ties them to week-one launch capacity
Is V4 proprietary?Midjourney describes it as new code + a new architecture designed internally
What trained V4?Cloud TPU v4 with JAX, according to Google

The roughly $320,000 Stable Diffusion 1.5 compute benchmark gives a possible scale for a substantial text-to-image training job of that era, nothing more.13

If somebody announces tomorrow that “Midjourney V1 cost exactly $100,000,” the useful first question is not whether $100,000 sounds plausible. It is whether they can show the invoice.

What Midjourney actually invented

It would be excessive to credit Midjourney with inventing diffusion models, CLIP or even the basic idea of steering visual generation with language. The papers, repositories and public experiments tell a much more collective story.5 6 7

It would be equally reductive to say the company merely packaged other people's work.

Its clearest early innovation may have been recognizing that diffusion could become a high-frequency social product, not just a research demonstration. In 2026 Holz claimed Midjourney was the first company to ship diffusion models at this scale. That is his claim and should remain attributed, but the scaling challenge described both by him and by Discord is real.1 16

Then came the preference loop. Midjourney did not ask users to occasionally rate a static benchmark. It placed choice inside the normal interface: four images, a selection, a variation, an upscale, another prompt. V3 shows the company was already turning those signals into model improvements within months of launch.2

Finally came the business model. Charging early enough for usage to pay for compute allowed Midjourney to turn customers into research financing. That is less glamorous than a new attention mechanism, but it may be just as important in explaining why the lab exists in its current form.

Holz often says he does not want to build an imaginative machine so much as expand the imagination of people.3 There is plenty to argue about in that framing, especially for artists whose work entered training data without consent. But the idea maps to a concrete product decision: Midjourney was built around exploration, selection and human reaction rather than a single “make final image” button.

In 2026, Holz described an internal ambition he calls an “aesthetic singularity”: collectively producing more visual exploration in a year than in all previous human history.1

That is grand enough to remain an ambition rather than a metric. But it helps explain the first Midjourney in retrospect.

The starting point was not a magical model sealed inside a basement. It was a set of newly published research breakthroughs, a small team able to assemble and extend them quickly, a founder willing to burn his own capital, and a Discord room where every person choosing among four images was also, often without thinking about it, teaching the system something about preference.

The model mattered. The loop mattered more.