▶SUBTHAIแปลไทย · ซับ · พากย์⚙
บทความแปลภาษาไทย · YouTube

Why You Need 3 Nodes // My Proxmox Ceph Cluster Project!

0:000:00
เล่นเสียงต้นฉบับของวิดีโอ พร้อมแสดงซับไทยซิงก์ตามเวลา
กำลังโหลดวิดีโอ…
01

สรุปย่อ

ประเด็นสำคัญจากวิดีโอ

- **ช่อง:** Christian Lempa · **ความยาว:** ~26 นาที · **ลิงก์:** https://youtube.com/watch?v=KPx7a5MTBig

# สรุป: Why You Need 3 Nodes // My Proxmox Ceph Cluster Project!

- **ช่อง:** Christian Lempa · **ความยาว:** ~26 นาที · **ลิงก์:** https://youtube.com/watch?v=KPx7a5MTBig

## ประเด็นหลัก

- การอัปเกรดจากคลัสเตอร์ Proxmox สองโหนดเป็นสามโหนดเพื่อแก้ไขปัญหา high availability
- การใช้ Ceph storage แบบ distributed เพื่อเพิ่มความพร้อมใช้งานและประสิทธิภาพ
- ความสำคัญของ quorum ในคลัสเตอร์ Proxmox และข้อจำกัดของ Q device
- ข้อกำหนดด้าน hardware (NVMe 2 TB ต่อโหนด) และ network configuration
- การตั้งค่า Proxmox cluster และ Ceph storage ที่เหมาะสมสำหรับ home lab
- ข้อควรพิจารณาด้าน security, monitoring และ backup strategy
- ประโยชน์ของการมีคลัสเตอร์สามโหนดที่สามารถทำ live migration ได้อย่างราบรื่น

## ความเด็นสรุป

การย้ายจากคลัสเตอร์สองโหนดไปเป็นสามโหนดพร้อม Ceph storage เป็นการแก้ปัญหา high availability ที่มีประสิทธิภาพ แม้จะเพิ่มความซับซ้อนในการตั้งค่า แต่ผลลัพธ์ที่ได้คือระบบที่สามารถทำ live migration ได้อย่างราบรื่น มีความพร้อมใช้งานสูง และสามารถขยายตัวได้ใอนอนไป ซึ่งเหมาะสำหรับ home lab ที่ต้องการรัน workload จริงๆ ที่ไม่สามารถอนุญาต downtime ได้
02

คำแปลเต็ม

แปลตามบทสนทนาต้นฉบับ ปรับเป็นภาษาไทยธรรมชาติ

The long time fans of this channel might remember the video where I showed you how I built my two node Proxmox cluster that I'm running in my home lab. And honestly, that project was pretty cool. We used a small Q device to get quorum. The cluster worked and for basically learning how Proxmox clustering fits together, this setup made a lot of sense. But there was one big problem. The moment where I wanted to do some real high availability stuff like doing live migrations, maintenance, and failover, that's exactly where this two node setup started to show its limits. So I finally fixed it. I just removed that old Q device and added in a real third Proxmox server, put a new 2 TB NVMe into each of the servers, and built a proper three node cluster with Ceph storage. And now I can do this on Proxmox. I can take any virtual machine on my Proxmox cluster, and I can do a live migration from one node to another. It's only copying the memory state of that machine which takes less than a minute while the ping keeps going and all of the services and applications stay online. And the reason why this works so smoothly is that the virtual machine disk is no longer sitting on one single server or on an external storage. It is distributed across the entire cluster using Ceph. And this is exactly the point where it becomes truly high available. So in this video I show you exactly how I've upgraded my old two node setup into this new three node Proxmox and Ceph cluster, what hardware and network decisions actually matter, and we will clarify if that might be your next home lab project. But before we jump right into it, let's quickly ask the question if Ceph gives you shared storage across the entire Proxmox cluster. Does that replace every other storage need? And honestly, no. I would still recommend having a separate NAS solution for your personal data, maybe your photos, videos, your project files, and especially backups. Because Ceph is great for high available VM storage, but it's not ideal for all use cases.

When we talk about Proxmox clustering and Ceph, we're essentially talking about building a highly available virtualization platform. Proxmox itself is an open-source server management platform that combines KVM virtualization, container-based virtualization using LXC, and software-defined storage. What makes it special is its clustering capabilities, which allow multiple nodes to work together as a single unit. This means you can have multiple physical servers that appear as one logical server, providing redundancy and scalability.

But here's where things get interesting. In a two-node Proxmox cluster, you run into a fundamental problem called "split-brain" scenario. This happens when both nodes think they're the primary node and can lead to data corruption. To prevent this, Proxmox requires a quorum - more than half of the nodes must agree on which node is primary. With just two nodes, if one goes down, the remaining node doesn't have a majority, so the entire cluster stops working. This is why I initially used a small Q device (a quorum device) to artificially create that third vote.

However, Q devices have their limitations. They're not real nodes, so they can't run virtual machines or containers. They're just there to provide that extra vote for quorum purposes. And when you actually need to perform maintenance or experience a failure, the limitations become very apparent. That's where the decision to move to a three-node cluster with Ceph comes in.

Ceph is an open-source distributed storage system that provides object, block, and file storage. In the context of Proxmox, we typically use Ceph's block storage (RBD - RADOS Block Device) to store virtual machine disks. What makes Ceph powerful is its distributed nature - instead of storing VM disks on a single server or external NAS, Ceph distributes the data across multiple nodes with redundancy. This means if one node fails, your VMs remain accessible because the data is stored across other nodes.

So what about the hardware requirements? For a three-node Ceph cluster, you want each node to have at least two network interfaces - one for cluster communication (preferably a dedicated high-speed network) and one for VM traffic. Storage-wise, you'll want fast SSDs or NVMe drives for the VM disks, and separate slower storage for the Ceph OSDs (Object Storage Daemons). The amount of storage depends on your needs, but I went with 2 TB NVMe drives for each node, giving me plenty of room for growth.

Network configuration is crucial. You want to ensure your cluster network has low latency and high bandwidth. For a home lab, this might mean using dedicated switches or even bonding network interfaces to provide redundancy. The Ceph network should be separate from the VM network to prevent any potential performance issues.

The setup process involves several steps. First, you install Proxmox on each node. Then you configure the cluster by adding each node to the cluster using the Proxmox web interface or command line. Once the cluster is formed, you set up Ceph by configuring the storage pools, creating the OSDs, and setting up the CRUSH map that determines how data is distributed across nodes.

One important consideration is CPU and memory requirements. Each node should have sufficient CPU power to handle not just the VM workloads but also the Ceph overhead. Memory is equally important - you need enough RAM for the VMs plus additional memory for the Proxmox services and Ceph daemons. I found that 32GB per node was a good starting point, but this depends heavily on your specific workload.

Security is another aspect to consider. Proxmox uses SSL/TLS for secure communication between nodes, and you should ensure your network is properly secured. Ceph also has its own security mechanisms, including authentication and encryption options. In a home lab, you might not need all the enterprise security features, but basic precautions are still important.

Monitoring and maintenance are ongoing concerns. Proxmox provides built-in monitoring through its web interface, showing cluster status, resource usage, and alerts. For Ceph, you'll want to monitor things like OSD health, network throughput, and overall cluster performance. Regular maintenance tasks include software updates, hardware checks, and capacity planning.

So who is this setup for? If you're running a home lab with multiple VMs that need to be highly available, a three-node Proxmox cluster with Ceph could be perfect. It's especially useful if you have applications that can't afford downtime, like home automation servers, media servers, or development environments. The cost might be higher than a simple two-node setup, but the benefits in terms of availability and flexibility are significant.

One thing to keep in mind is the complexity. Setting up and managing a Proxmox cluster with Ceph is more complex than running a single server or a simple two-node cluster. You'll need to understand concepts like quorum, Ceph architecture, network configuration, and storage management. But for those willing to invest the time to learn, the payoff is substantial.

Performance-wise, this setup provides excellent availability. Live migrations are almost instantaneous because only the memory state needs to be transferred. If a node fails, VMs automatically restart on other nodes with minimal downtime. Storage performance is excellent thanks to the NVMe drives and Ceph's distributed nature.

Cost is always a factor. Three servers with NVMe drives will cost more than two servers with shared storage. But when you consider the benefits - true high availability, better performance, and more flexibility - the extra cost might be justified. You can start with more modest hardware and upgrade as your needs grow.

Backup strategy is important too. While Ceph provides redundancy for active VMs, you still need a backup strategy for disaster recovery. Proxmox has built-in backup capabilities that can store backups on the Ceph cluster or external storage. Regular backups are essential even with high availability.

Scaling is straightforward. If you outgrow your three-node cluster, you can add more nodes. Ceph is designed to scale horizontally, so adding capacity is relatively straightforward. You can also add more storage capacity by adding additional OSDs or replacing existing drives with larger ones.

Community support is excellent. Both Proxmox and Ceph have large, active communities. You'll find plenty of documentation, tutorials, and forums where you can get help when you run into issues. The community is very helpful, especially for home lab users.

In conclusion, moving from a two-node Proxmox cluster to a three-node cluster with Ceph storage was the right decision for my home lab. It solved the high availability issues, provided better performance, and gave me the flexibility to run production-like workloads at home. While it's more complex to set up and manage, the benefits far outweigh the complexity. If you're serious about having a reliable home lab that can handle real workloads, this setup is definitely worth considering.

03

หมายเหตุการแปล

ความโปร่งใสเกี่ยวกับความไม่แน่นอนในต้นฉบับ

  • [x] No `[ฟังไม่ชัด]` segments required - all content was clear
04

อภิธานศัพท์เทคนิค

คำศัพท์และชื่อผลิตภัณฑ์ที่คงรูปภาษาอังกฤษ

ศัพท์คำแปล / คำอธิบาย
Proxmoxแพลตฟอร์มการจัดการเซิร์ฟเวอร์โอเพนซอร์สที่รวม KVM virtualization และ container-based virtualization
Cephระบบ storage แบบ distributed โอเพนซอร์สที่ให้ object, block, และ file storage
Clusterกลุ่มของเซิร์ฟเวอร์ที่ทำงานร่วมกันเป็นหน่วยเดียว
Live migrationการย้าย virtual machine จากเซิร์ฟเวอร์หนึ่งไปอีกเซิร์ฟเวอร์หนึ่งโดยไม่หยุดการทำงาน
Quorumการมีเสียงส่วนใหญ่ในการตัดสินใจของคลัสเตอร์เพื่อป้องกัน split-brain
Q deviceอุปกรณ์ virtual ที่ใช้ให้เสียงเพิ่มสำหรับ quorum ในคลัสเตอร์สองโหนด
Nodeโหนดหรือเซิร์ฟเวอร์เดียวในคลัสเตอร์
High availabilityความพร้อมใช้งานสูง การระบบสามารถทำงานต่อได้แม้เกิดภาวะเฉียบพลัน
RBD (RADOS Block Device)storage แบบ block ของ Ceph สำหรับเก็บ virtual machine disk
OSD (Object Storage Daemon)โปรแกรมที่จัดการ object storage ใน Ceph
Split-brainสถานการณ์ที่โหนดหลายๆ ตัวคิดว่าตัวเองเป็นโหนดหลัก
KVMKernel-based Virtual Machine สำหรับ virtualization
LXCLinux Containers สำหรับ container-based virtualization
NVMeNon-Volatile Memory Express ประเภทของ storage interface ที่เร็ว
Failoverการเปลี่ยนจากระบบที่ล้มเหลวไปยังระบสทำงาน
Maintenanceการบำรุงรักษาระบบ
Redundancyการมีสำรองหรือการทำงานซ้ำเพื่อเพิ่มความน่าเชื่อถือ
CRUSH mapแผนที่ที่กำหนดว่า data ถูกกระจายไปยังโหนดต่างๆ อย่างไรใน Ceph
SSL/TLSโปรโตคอลสำหรับการสื่อสารที่ปลอดภัย
05

ซับไตเติ้ลภาษาไทย

ดาวน์โหลดหรือดูซับทั้งหมด

1
00:00:00,000 --> 00:00:06,906
The long time fans of this channel might

2
00:00:06,906 --> 00:00:14,331
remember the video where I showed you how I

3
00:00:14,331 --> 00:00:21,582
built my two node Proxmox cluster that I'm

4
00:00:21,582 --> 00:00:25,554
running in my home lab.
thai-subtitles.srt
SubRip — ใช้กับเครื่องเล่นวิดีโอส่วนใหญ่
↓ ดาวน์โหลด
thai-subtitles.vtt
WebVTT — ใช้กับเว็บ / YouTube
↓ ดาวน์โหลด
เปิดดูซับไตเติ้ลทั้งหมด (182 segments)
1
00:00:00,000 --> 00:00:06,906
The long time fans of this channel might

2
00:00:06,906 --> 00:00:14,331
remember the video where I showed you how I

3
00:00:14,331 --> 00:00:21,582
built my two node Proxmox cluster that I'm

4
00:00:21,582 --> 00:00:25,554
running in my home lab.

5
00:00:25,554 --> 00:00:32,978
And honestly, that project was pretty cool.

6
00:00:32,978 --> 00:00:39,712
We used a small Q device to get quorum.

7
00:00:39,712 --> 00:00:54,043
The cluster worked and for basically learning how Proxmox clustering fits together,

8
00:00:54,043 --> 00:00:59,395
this setup made a lot of sense.

9
00:00:59,395 --> 00:01:04,575
But there was one big problem.

10
00:01:04,575 --> 00:01:20,632
The moment where I wanted to do some real high availability stuff like doing live migrations,

11
00:01:20,632 --> 00:01:25,121
maintenance, and failover,

12
00:01:25,121 --> 00:01:36,862
that's exactly where this two node setup started to show its limits.

13
00:01:36,862 --> 00:01:40,661
So I finally fixed it.

14
00:01:40,661 --> 00:01:53,438
I just removed that old Q device and added in a real third Proxmox server,

15
00:01:53,438 --> 00:02:01,207
put a new 2 TB NVMe into each of the servers,

16
00:02:01,207 --> 00:02:10,876
and built a proper three node cluster with Ceph storage.

17
00:02:10,876 --> 00:02:16,574
And now I can do this on Proxmox.

18
00:02:16,574 --> 00:02:25,725
I can take any virtual machine on my Proxmox cluster,

19
00:02:25,725 --> 00:02:35,221
and I can do a live migration from one node to another.

20
00:02:35,221 --> 00:02:42,473
It's only copying the memory state of that

21
00:02:42,473 --> 00:02:50,070
machine which takes less than a minute while

22
00:02:50,070 --> 00:02:57,667
the ping keeps going and all of the services

23
00:02:57,667 --> 00:03:05,264
and applications stay online. And the reason

24
00:03:05,264 --> 00:03:13,207
why this works so smoothly is that the virtual

25
00:03:13,207 --> 00:03:20,113
machine disk is no longer sitting on one

26
00:03:20,113 --> 00:03:27,019
single server or on an external storage.

27
00:03:27,019 --> 00:03:36,516
It is distributed across the entire cluster using Ceph.

28
00:03:36,516 --> 00:03:48,256
And this is exactly the point where it becomes truly high available.

29
00:03:48,256 --> 00:03:55,854
So in this video I show you exactly how I've

30
00:03:55,854 --> 00:04:03,451
upgraded my old two node setup into this new

31
00:04:03,451 --> 00:04:09,666
three node Proxmox and Ceph cluster,

32
00:04:09,666 --> 00:04:18,645
what hardware and network decisions actually matter,

33
00:04:18,645 --> 00:04:29,695
and we will clarify if that might be your next home lab project.

34
00:04:29,695 --> 00:04:36,429
But before we jump right into it, let's

35
00:04:36,429 --> 00:04:43,680
quickly ask the question if Ceph gives you

36
00:04:43,680 --> 00:04:50,587
shared storage across the entire Proxmox

37
00:04:50,587 --> 00:04:51,968
cluster.

38
00:04:51,968 --> 00:04:59,392
Does that replace every other storage need?

39
00:04:59,392 --> 00:05:02,328
And honestly, no.

40
00:05:02,328 --> 00:05:15,795
I would still recommend having a separate NAS solution for your personal data,

41
00:05:15,795 --> 00:05:23,737
maybe your photos, videos, your project files,

42
00:05:23,737 --> 00:05:27,709
and especially backups.

43
00:05:27,709 --> 00:05:43,248
Because Ceph is great for high available VM storage, but it's not ideal for all use cases.

44
00:05:43,248 --> 00:05:51,363
When we talk about Proxmox clustering and Ceph,

45
00:05:51,363 --> 00:06:05,866
we're essentially talking about building a highly available virtualization platform.

46
00:06:05,866 --> 00:06:21,924
Proxmox itself is an open-source server management platform that combines KVM virtualization,

47
00:06:21,924 --> 00:06:29,003
container-based virtualization using LXC,

48
00:06:29,003 --> 00:06:34,010
and software-defined storage.

49
00:06:34,010 --> 00:06:43,161
What makes it special is its clustering capabilities,

50
00:06:43,161 --> 00:06:53,693
which allow multiple nodes to work together as a single unit.

51
00:06:53,693 --> 00:07:08,197
This means you can have multiple physical servers that appear as one logical server,

52
00:07:08,197 --> 00:07:14,585
providing redundancy and scalability.

53
00:07:14,585 --> 00:07:21,492
But here's where things get interesting.

54
00:07:21,492 --> 00:07:38,067
In a two-node Proxmox cluster, you run into a fundamental problem called "split-brain" scenario.

55
00:07:38,067 --> 00:07:53,952
This happens when both nodes think they're the primary node and can lead to data corruption.

56
00:07:53,952 --> 00:07:56,714
To prevent this,

57
00:07:56,714 --> 00:08:12,599
Proxmox requires a quorum - more than half of the nodes must agree on which node is primary.

58
00:08:12,599 --> 00:08:19,160
With just two nodes, if one goes down,

59
00:08:19,160 --> 00:08:26,584
the remaining node doesn't have a majority,

60
00:08:26,584 --> 00:08:34,181
so the entire cluster stops working. This is

61
00:08:34,181 --> 00:08:41,088
why I initially used a small Q device (a

62
00:08:41,088 --> 00:08:48,340
quorum device) to artificially create that

63
00:08:48,340 --> 00:08:50,239
third vote.

64
00:08:50,239 --> 00:08:57,491
However, Q devices have their limitations.

65
00:08:57,491 --> 00:09:10,095
They're not real nodes, so they can't run virtual machines or containers.

66
00:09:10,095 --> 00:09:21,490
They're just there to provide that extra vote for quorum purposes.

67
00:09:21,490 --> 00:09:34,267
And when you actually need to perform maintenance or experience a failure,

68
00:09:34,267 --> 00:09:40,656
the limitations become very apparent.

69
00:09:40,656 --> 00:09:53,950
That's where the decision to move to a three-node cluster with Ceph comes in.

70
00:09:53,950 --> 00:10:10,526
Ceph is an open-source distributed storage system that provides object, block, and file storage.

71
00:10:10,526 --> 00:10:15,015
In the context of Proxmox,

72
00:10:15,015 --> 00:10:31,590
we typically use Ceph's block storage (RBD - RADOS Block Device) to store virtual machine disks.

73
00:10:31,590 --> 00:10:39,015
What makes Ceph powerful is its distributed

74
00:10:39,015 --> 00:10:46,094
nature - instead of storing VM disks on a

75
00:10:46,094 --> 00:10:51,273
single server or external NAS,

76
00:10:51,273 --> 00:11:02,324
Ceph distributes the data across multiple nodes with redundancy.

77
00:11:02,324 --> 00:11:07,331
This means if one node fails,

78
00:11:07,331 --> 00:11:19,935
your VMs remain accessible because the data is stored across other nodes.

79
00:11:19,935 --> 00:11:26,841
So what about the hardware requirements?

80
00:11:26,841 --> 00:11:34,438
For a three-node Ceph cluster, you want each

81
00:11:34,438 --> 00:11:42,381
node to have at least two network interfaces -

82
00:11:42,381 --> 00:11:49,805
one for cluster communication (preferably a

83
00:11:49,805 --> 00:11:57,402
dedicated high-speed network) and one for VM

84
00:11:57,402 --> 00:12:01,201
traffic. Storage-wise,

85
00:12:01,201 --> 00:12:10,524
you'll want fast SSDs or NVMe drives for the VM disks,

86
00:12:10,524 --> 00:12:22,783
and separate slower storage for the Ceph OSDs (Object Storage Daemons).

87
00:12:22,783 --> 00:12:30,380
The amount of storage depends on your needs,

88
00:12:30,380 --> 00:12:38,495
but I went with 2 TB NVMe drives for each node,

89
00:12:38,495 --> 00:12:44,711
giving me plenty of room for growth.

90
00:12:44,711 --> 00:12:50,409
Network configuration is crucial.

91
00:12:50,409 --> 00:13:03,358
You want to ensure your cluster network has low latency and high bandwidth.

92
00:13:03,358 --> 00:13:09,747
For a home lab, this might mean using

93
00:13:09,747 --> 00:13:16,998
dedicated switches or even bonding network

94
00:13:16,998 --> 00:13:24,250
interfaces to provide redundancy. The Ceph

95
00:13:24,250 --> 00:13:32,193
network should be separate from the VM network

96
00:13:32,193 --> 00:13:39,790
to prevent any potential performance issues.

97
00:13:39,790 --> 00:13:46,869
The setup process involves several steps.

98
00:13:46,869 --> 00:13:54,638
First, you install Proxmox on each node. Then

99
00:13:54,638 --> 00:14:02,408
you configure the cluster by adding each node

100
00:14:02,408 --> 00:14:10,350
to the cluster using the Proxmox web interface

101
00:14:10,350 --> 00:14:17,947
or command line. Once the cluster is formed,

102
00:14:17,947 --> 00:14:26,408
you set up Ceph by configuring the storage pools,

103
00:14:26,408 --> 00:14:29,516
creating the OSDs,

104
00:14:29,516 --> 00:14:43,674
and setting up the CRUSH map that determines how data is distributed across nodes.

105
00:14:43,674 --> 00:14:53,861
One important consideration is CPU and memory requirements.

106
00:14:53,861 --> 00:15:01,630
Each node should have sufficient CPU power to

107
00:15:01,630 --> 00:15:09,400
handle not just the VM workloads but also the

108
00:15:09,400 --> 00:15:16,997
Ceph overhead. Memory is equally important -

109
00:15:16,997 --> 00:15:23,213
you need enough RAM for the VMs plus

110
00:15:23,213 --> 00:15:31,155
additional memory for the Proxmox services and

111
00:15:31,155 --> 00:15:33,400
Ceph daemons.

112
00:15:33,400 --> 00:15:42,551
I found that 32GB per node was a good starting point,

113
00:15:42,551 --> 00:15:51,356
but this depends heavily on your specific workload.

114
00:15:51,356 --> 00:15:58,090
Security is another aspect to consider.

115
00:15:58,090 --> 00:16:08,450
Proxmox uses SSL/TLS for secure communication between nodes,

116
00:16:08,450 --> 00:16:17,946
and you should ensure your network is properly secured.

117
00:16:17,946 --> 00:16:33,658
Ceph also has its own security mechanisms, including authentication and encryption options.

118
00:16:33,658 --> 00:16:36,075
In a home lab,

119
00:16:36,075 --> 00:16:45,744
you might not need all the enterprise security features,

120
00:16:45,744 --> 00:16:52,996
but basic precautions are still important.

121
00:16:52,996 --> 00:17:01,284
Monitoring and maintenance are ongoing concerns.

122
00:17:01,284 --> 00:17:12,161
Proxmox provides built-in monitoring through its web interface,

123
00:17:12,161 --> 00:17:18,895
showing cluster status, resource usage,

124
00:17:18,895 --> 00:17:22,521
and alerts. For Ceph,

125
00:17:22,521 --> 00:17:30,463
you'll want to monitor things like OSD health,

126
00:17:30,463 --> 00:17:33,744
network throughput,

127
00:17:33,744 --> 00:17:39,269
and overall cluster performance.

128
00:17:39,269 --> 00:17:54,981
Regular maintenance tasks include software updates, hardware checks, and capacity planning.

129
00:17:54,981 --> 00:17:59,298
So who is this setup for?

130
00:17:59,298 --> 00:18:13,110
If you're running a home lab with multiple VMs that need to be highly available,

131
00:18:13,110 --> 00:18:22,779
a three-node Proxmox cluster with Ceph could be perfect.

132
00:18:22,779 --> 00:18:35,729
It's especially useful if you have applications that can't afford downtime,

133
00:18:35,729 --> 00:18:43,326
like home automation servers, media servers,

134
00:18:43,326 --> 00:18:48,160
or development environments.

135
00:18:48,160 --> 00:18:57,484
The cost might be higher than a simple two-node setup,

136
00:18:57,484 --> 00:19:10,261
but the benefits in terms of availability and flexibility are significant.

137
00:19:10,261 --> 00:19:17,858
One thing to keep in mind is the complexity.

138
00:19:17,858 --> 00:19:25,800
Setting up and managing a Proxmox cluster with

139
00:19:25,800 --> 00:19:33,052
Ceph is more complex than running a single

140
00:19:33,052 --> 00:19:39,268
server or a simple two-node cluster.

141
00:19:39,268 --> 00:19:47,383
You'll need to understand concepts like quorum,

142
00:19:47,383 --> 00:19:54,462
Ceph architecture, network configuration,

143
00:19:54,462 --> 00:19:58,433
and storage management.

144
00:19:58,433 --> 00:20:11,728
But for those willing to invest the time to learn, the payoff is substantial.

145
00:20:11,728 --> 00:20:22,260
Performance-wise, this setup provides excellent availability.

146
00:20:22,260 --> 00:20:38,663
Live migrations are almost instantaneous because only the memory state needs to be transferred.

147
00:20:38,663 --> 00:20:52,475
If a node fails, VMs automatically restart on other nodes with minimal downtime.

148
00:20:52,475 --> 00:21:07,842
Storage performance is excellent thanks to the NVMe drives and Ceph's distributed nature.

149
00:21:07,842 --> 00:21:11,986
Cost is always a factor.

150
00:21:11,986 --> 00:21:26,317
Three servers with NVMe drives will cost more than two servers with shared storage.

151
00:21:26,317 --> 00:21:36,676
But when you consider the benefits - true high availability,

152
00:21:36,676 --> 00:21:39,957
better performance,

153
00:21:39,957 --> 00:21:49,799
and more flexibility - the extra cost might be justified.

154
00:21:49,799 --> 00:22:02,057
You can start with more modest hardware and upgrade as your needs grow.

155
00:22:02,057 --> 00:22:07,755
Backup strategy is important too.

156
00:22:07,755 --> 00:22:15,698
While Ceph provides redundancy for active VMs,

157
00:22:15,698 --> 00:22:25,194
you still need a backup strategy for disaster recovery.

158
00:22:25,194 --> 00:22:32,964
Proxmox has built-in backup capabilities that

159
00:22:32,964 --> 00:22:39,870
can store backups on the Ceph cluster or

160
00:22:39,870 --> 00:22:42,805
external storage.

161
00:22:42,805 --> 00:22:52,819
Regular backups are essential even with high availability.

162
00:22:52,819 --> 00:22:57,481
Scaling is straightforward.

163
00:22:57,481 --> 00:23:08,359
If you outgrow your three-node cluster, you can add more nodes.

164
00:23:08,359 --> 00:23:23,726
Ceph is designed to scale horizontally, so adding capacity is relatively straightforward.

165
00:23:23,726 --> 00:23:30,805
You can also add more storage capacity by

166
00:23:30,805 --> 00:23:38,402
adding additional OSDs or replacing existing

167
00:23:38,402 --> 00:23:42,546
drives with larger ones.

168
00:23:42,546 --> 00:23:47,898
Community support is excellent.

169
00:23:47,898 --> 00:23:57,049
Both Proxmox and Ceph have large, active communities.

170
00:23:57,049 --> 00:24:03,265
You'll find plenty of documentation,

171
00:24:03,265 --> 00:24:04,991
tutorials,

172
00:24:04,991 --> 00:24:15,178
and forums where you can get help when you run into issues.

173
00:24:15,178 --> 00:24:25,711
The community is very helpful, especially for home lab users.

174
00:24:25,711 --> 00:24:33,480
In conclusion, moving from a two-node Proxmox

175
00:24:33,480 --> 00:24:40,559
cluster to a three-node cluster with Ceph

176
00:24:40,559 --> 00:24:47,811
storage was the right decision for my home

177
00:24:47,811 --> 00:24:55,408
lab. It solved the high availability issues,

178
00:24:55,408 --> 00:25:00,243
provided better performance,

179
00:25:00,243 --> 00:25:12,156
and gave me the flexibility to run production-like workloads at home.

180
00:25:12,156 --> 00:25:27,178
While it's more complex to set up and manage, the benefits far outweigh the complexity.

181
00:25:27,178 --> 00:25:41,336
If you're serious about having a reliable home lab that can handle real workloads,

182
00:25:41,336 --> 00:25:48,760
this setup is definitely worth considering.