Minisforum MS-S1 MAX-P495 Cluster Price

Minisforum MS-S1 MAX-P495 Cluster Price

A few days ago, the strangest part of Minisforum’s new AI workstation was the hardware.

One MS-S1 MAX-P495 has 192GB of LPDDR5X-8533 unified memory. Four of them give you 768GB of aggregate memory. Minisforum says a four-unit cluster successfully ran a DeepSeek-R1 671B Q4_0 model weighing about 380GB.

At the time, there was one missing number.

The price.

Now we have it.

Minisforum’s US store lists the 192GB + 2TB MS-S1 MAX-P495 at $9,249, with a current $7,399 limited-time price.

Multiply the normal listed price by four: $36,996.

Even at the discounted price: $29,596.

So yes, four compact PCs can apparently hold and run one of the biggest open-weight models people talk about.

But the moment you put a price on the cluster, this stops feeling like a clever mini-PC trick.

It starts feeling like a $37,000 AI server built out of four small boxes.

And that makes the comparison much more interesting.

Access without medium partner: Four Mini PCs Run DeepSeek-R1 671B

Generated By Author

The price finally answers the question the spec sheet could not

I liked the original MS-S1 MAX-P495 story because the numbers were absurd.

192GB per machine.

160GB allocatable as VRAM.

Two 10GbE ports.

120W sustained APU power.

2U rack support.

Four systems running a 671B model.

It looked like somebody took the local-AI mini-PC trend and kept turning the memory dial until it became a server.

But without a price, it was hard to judge.

A four-node cluster for $12,000 would be one kind of product.

A four-node cluster for $20,000 would be another.

At nearly $37,000 list price, we can finally stop treating this like a cheap workaround.

It is compact infrastructure.

That does not make it bad.

It just changes the entire conversation.

One machine costs more than many full workstations

The current official US product page shows:

$9,249 list price.

$7,399 current sale price.

That gets you the Ryzen AI Max+ PRO 495, 192GB LPDDR5X-8533, and a 2TB SSD.

The processor itself is interesting.

It has 16 Zen 5 CPU cores and 32 threads, Radeon 8065S integrated graphics, an NPU rated at up to 55 TOPS, and Minisforum claims up to 131 TOPS overall AI compute for the system.

The GPU can use as much as 160GB of the shared memory as VRAM.

That is the reason this box exists.

You are not paying nine thousand dollars because you need a small Windows PC.

You are paying for an unusual memory architecture that lets one compact machine hold models far larger than a normal consumer GPU can.

An RTX 5090 has 32GB of VRAM.

This box can give its GPU up to 160GB.

The 5090 will be much faster on many workloads that fit inside 32GB.

The Minisforum box wins a different game:

Can the model fit at all?

Four machines means 768GB. But not one 768GB GPU.

This is where the giant number needs a warning label.

Four times 192GB is obviously 768GB.

That math is real.

But the memory lives inside four separate computers.

Each machine has its own processor.

Its own memory.

Its own GPU.

Its own operating environment.

The model has to be split across nodes.

Data has to move between them.

The nodes have to wait for each other.

That means 768GB aggregate memory is not the same thing as 768GB of local GPU memory.

If somebody puts “768GB AI cluster” in a headline, fine.

If they call it a “768GB GPU,” no.

The network is now part of the inference engine.

That changes latency and throughput in ways the memory number alone cannot show.

Minisforum says DeepSeek-R1 671B ran. It does not tell us how fast.

The official product page makes a very specific claim.

Minisforum says a four-unit cluster successfully ran DeepSeek-R1 671B Q4, using a roughly 380GB Q4_0 model.

That is enough to prove one thing:

The model can be distributed across the cluster and loaded.

What Minisforum does not publish on that page is a token-per-second result for the four-node DeepSeek run.

That missing number matters.

A model loading is not the same as a model being pleasant to use.

If it outputs 1.8 tokens per second, technically the 671B model runs.

If it outputs 15 or 20 tokens per second, we are talking about a very different machine.

This is the benchmark I want next.

Not another “supported model size” badge.

Show me prompt processing speed, output tokens per second, time to first token, power draw, and speed after an hour of continuous inference.

That would tell us what the $37,000 actually buys.

The two-node result is more useful right now

Minisforum gives us a better number for the smaller cluster.

Two 192GB units give you 384GB aggregate memory.

The company says that pair can run Qwen3.5–397B at 16 tok/s.

That is far more useful than “it loads.”

Sixteen tokens per second is a speed a human can comfortably read.

Again, it is a vendor benchmark. I want independent testing.

But it tells us the system is not only a memory stunt.

And the cost is easier to swallow, relatively speaking.

Two machines at list price:

$18,498.

Two at the current discounted price:

$14,798.

Still expensive.

But if your actual workload fits into two nodes, the difference between $14,798 and $29,596 is not small.

This is why I would not build a four-node cluster simply because four looks good in a rack.

Buy the number of nodes your model actually needs.

Nearly $30,000 on sale is still nearly $30,000

The discount deserves context because the headline uses the normal listed price.

Minisforum is currently taking 20% off, dropping each machine from $9,249 to $7,399.

That saves $1,850 per machine.

Across four machines, the saving is $7,400.

Which sounds great until you look at the checkout total.

$29,596.

Before whatever extra rack hardware, networking, storage expansion or backup equipment you decide to add.

The sale price does not turn this into a bargain.

It turns a $37,000 cluster into a $30,000 cluster.

That is a meaningful difference.

It is also still serious money.

And the cluster needs more than four boxes

This is another place where product-page arithmetic gets too clean.

You cannot only buy four PCs and call the project finished.

If you want a proper rack setup, you need the physical rack arrangement.

You need networking.

You may want a switch depending on topology.

You need cables.

You probably want a UPS if this is being used for actual work.

You need enough power delivery.

Then storage becomes a question too.

Each P495 comes with 2TB in the current configuration, but large model collections get big very fast. DeepSeek-R1 Q4 alone is around 380GB in the example Minisforum cites.

A few model variants, checkpoints and datasets later, 2TB does not look enormous.

The machine does support two M.2 slots and RAID.

That helps.

It also means the real cluster bill can move beyond the simple four-times-$9,249 calculation.

Four nodes can pull serious power, even if they are small

The boxes look compact.

The power budget is not toy-like.

Minisforum says the P495 supports 120W sustained PPT and 160W peak PPT.

Four units at the sustained setting means roughly 480W of APU power alone.

At peak, the four processors could theoretically total 640W, before PSU losses and other system components.

That is still far below some large multi-GPU server configurations.

But this is not a Raspberry Pi cluster.

Four of these machines running a 671B model for hours are real compute hardware.

You need cooling.

You need airflow.

And if electricity is expensive where you live, you will notice.

The small chassis can hide the scale of the workload.

The power meter does not care how cute the boxes are.

The dual 10GbE ports suddenly matter a lot more at $37,000

Each MS-S1 MAX-P495 includes dual 10GbE, plus USB4 V2 connectivity.

That is unusually serious networking for a compact workstation.

It also tells you what Minisforum had in mind.

This thing was never only a desk PC.

The company even includes a Rack mode, supports horizontal deployment, and provides a cascade power-on header so several machines can be started or stopped together.

Nice.

But I still want Minisforum to document the exact cluster setup used for the DeepSeek benchmark.

Which links carried model traffic?

What software split the model?

Was the model divided by layers?

What was communication overhead?

How much throughput disappeared as the node count increased from two to four?

Without those details, we know the model ran.

We do not know whether this was the smartest way to run it.

This is where NVIDIA becomes the obvious comparison

Once a local-AI box costs more than $9,000, the comparison pool gets much larger.

You are no longer deciding between a $900 mini PC and a gaming desktop.

You start looking at NVIDIA’s compact Grace Blackwell systems, GPU workstations, used enterprise cards and proper servers.

NVIDIA’s DGX Spark class is especially relevant because it attacks the same broad idea from the opposite direction.

Spark gives you 128GB of coherent unified memory, Grace Blackwell, NVIDIA’s CUDA stack and high-speed networking.

Minisforum gives you 192GB per node, x86 Zen 5, Radeon 8065S graphics and far more memory capacity once several machines are combined.

This is not an easy “more GB wins” comparison.

NVIDIA has much stronger AI software support.

The AMD system has a massive memory-capacity argument.

At $37,000, software friction starts costing real money too.

If one platform takes an engineer two extra days to make a model work properly, that is part of the price whether the invoice shows it or not.

The $37,000 number makes “local is cheaper than cloud” harder to say casually

One of the most common arguments for local AI is cost.

Buy the hardware once.

Stop paying API bills.

Run as much as you want.

Keep your data private.

That argument can be correct.

But a $29,596 to $36,996 cluster changes the break-even math.

If you are a company running a private model every day, the hardware may still make sense.

If the workload is sensitive and cannot leave the office, cloud cost may not even be the main concern.

If you need predictable capacity with no per-token fee, local hardware has value.

But for somebody experimenting with a 671B model a few hours per week?

Thirty-seven thousand dollars buys a lot of API usage.

And cloud GPUs can scale up when you need them and disappear when you do not.

Local AI is not automatically cheaper.

It depends on utilization.

That boring finance word matters more at this price.

Privacy may be the best argument for this cluster

There is one case where the price becomes easier to understand.

Private data.

Imagine a legal team that cannot send case material to an external provider.

Or a research lab working with unpublished data.

Or a company indexing its entire internal codebase.

A 671B-class local model may not be necessary for all those jobs, but keeping inference physically inside the office has value.

The cluster also avoids internet dependency.

Your token bill does not change because somebody at an API provider changed pricing.

You know where the data is.

You control when the system is patched.

For a business, that can matter more than whether a cloud GPU is cheaper per hour.

For a home user?

Different story.

I would not spend $30,000 because I dislike subscriptions.

The mini-PC part is almost becoming meaningless

This is the funniest thing about the MS-S1 MAX-P495.

Calling it a mini PC gives people the wrong mental image.

A normal mini PC is something you put behind a monitor.

Maybe it runs Plex.

Maybe it is a small development box.

Maybe it replaces an office tower.

This machine has 192GB of unified memory, dual 10GbE, PCIe expansion, 120W sustained APU power, rack deployment, a 320W internal PSU, and explicit multi-node controls.

Then Minisforum shows four of them running a 671B model.

That is not really the same category as a $500 NUC-style machine anymore.

“Mini workstation” is better.

“Compact AI node” is probably more accurate.

Four of them?

That is a server cluster wearing mini-PC clothes.

The memory capacity is still kind of absurd

I do not want the price criticism to erase what the hardware is doing.

One node has 192GB.

Four have 768GB aggregate.

That is a huge amount of model capacity in a small physical space.

An RTX 5090 has 32GB.

You would need an unrealistic number of consumer GPUs to match the memory capacity, and then power, PCIe lanes, chassis space and multi-GPU software become their own headache.

Enterprise GPUs solve those problems better.

They also cost enterprise money.

This is why the P495 exists.

AMD’s Ryzen AI Max platform occupies a strange middle area between consumer PCs and server accelerators.

It is slower than serious data-center GPUs.

It has much more memory than consumer GPUs.

And OEMs can put it into compact boxes.

For model capacity, that combination is hard to ignore.

But 768GB is more memory than DeepSeek-R1 Q4 needs

Here is another detail worth noticing.

Minisforum says its DeepSeek-R1 671B Q4_0 model was around 380GB.

The four-node cluster has 768GB total memory.

So the model itself does not need all 768GB just for weights.

Of course there is overhead.

You need KV cache.

You need runtime memory.

The operating systems need memory.

Distributed execution needs buffers.

Long context can consume a lot.

Still, the four-node setup gives the model a lot of headroom.

That raises a useful question:

Could a differently configured three-node setup handle the same model?

Three P495 systems would provide 576GB of aggregate memory.

At list pricing, that would cost $27,747.

At the current sale price:

$22,197.

Minisforum highlights two-node and four-node configurations, not a three-node DeepSeek benchmark.

I would love to see one.

If three nodes can run the same model at useful speed, suddenly one quarter of the cluster bill disappears.

The older 128GB MS-S1 makes the price jump look even stranger

Minisforum also sells the previous MS-S1 MAX 128GB Max AI Compute Edition.

The current official US store lists that machine around $3,799, discounted from $4,749.

The new P495 with 192GB is $7,399 on sale.

That is almost double the discounted price.

Yes, you get the newer PRO 495.

Yes, you get another 64GB of faster memory.

There are platform improvements.

But this tells us how expensive the jump to 192GB has become in 2026.

Memory is no longer a side spec.

It is a huge part of the product’s price.

The AI memory shortage is showing up directly inside these machines.

I would rather see a 397B benchmark than another 671B loading screenshot

This may sound backwards.

The 671B model gets clicks.

The 397B at 16 tok/s result is more useful.

Why?

Because it gives us an actual speed.

It tells us what the machine feels like.

If two nodes can consistently deliver 16 tok/s on a 397B model, that could be useful for serious local work.

The four-node 671B test needs the same treatment.

Give us the number.

Maybe it is 12 tok/s.

Maybe it is 5.

Maybe it is 1.5.

Those are completely different products from the user’s perspective.

The next Minisforum chart I want has fewer parameter counts and more tokens per second.

Nearly $37,000 sounds ridiculous until you compare it with enterprise AI hardware

This is where I have to pull back a little.

Thirty-seven thousand dollars is absurd for four “mini PCs.”

Thirty-seven thousand dollars is not absurd for AI infrastructure.

Enterprise accelerator systems can move past that number very quickly.

Professional GPUs with large HBM pools are expensive.

Server platforms need CPUs, memory, networking, storage and support contracts.

A serious GPU server can make $37,000 look normal.

That is why I do not think Minisforum has necessarily priced this product incorrectly.

The mistake would be comparing it only with consumer PCs.

The P495 is drifting into a different market.

The question becomes whether its performance per dollar, performance per watt, and software experience make sense against workstation and server alternatives.

We cannot answer that from model-size claims alone.

The sale price is probably the more important number for early buyers

If somebody is actually considering this system today, the practical cluster price is $29,596, not $36,996.

That is what four units cost at the current $7,399 price.

Minisforum says the discount is limited.

The US listing also says the 192GB + 2TB system has an estimated shipping time of mid-October.

So this is not theoretical anymore.

You can put four in a cart.

The price problem has gone from “we’ll see” to a number on a checkout page.

I think that makes this story much stronger than the launch announcement.

The hardware did not change.

Our understanding of what the hardware costs did.

I still like the machine. I like the cluster less.

One P495 makes sense to me.

192GB unified memory in a compact workstation is unusual enough to create real use cases.

Two make sense for people who genuinely need models beyond one node and can use the 384GB aggregate memory.

Four?

Now I want a business case.

At $29,596 on sale, the cluster needs to save real cloud money, solve a privacy problem, or enable a workload you cannot easily run another way.

“DeepSeek-R1 671B fits” is not enough.

It needs to run well.

It needs to stay stable.

The distributed software needs to be manageable.

The network overhead needs to be acceptable.

And somebody needs to maintain four machines instead of one.

That is the difference between a cool demo and infrastructure.

The missing number is no longer the price

Last week, I ended the conversation with a simple question:

How much?

Minisforum has answered it.

$9,249 per node.

$36,996 for four at list price.

$29,596 for four at today’s discounted price.

We know the memory.

We know the model size.

We know the rack format.

Now there is one number left that matters more than all of them:

How many tokens per second does DeepSeek-R1 671B actually produce on the four-node cluster?

Because at $37,000, “it runs” is no longer enough.

Sources I checked

Minisforum MS-S1 MAX-P495 official product page

Minisforum US store

TechRadar: Minisforum MS-S1 MAX-P495 pricing and preorder details

NVIDIA Personal AI Supercomputers marketplace


Post a Comment

Previous Post Next Post