Why You Need 3 Nodes // My Proxmox Ceph Cluster Project!
สรุปย่อ
ประเด็นสำคัญจากวิดีโอ
- **ช่อง:** Christian Lempa · **ความยาว:** ~26 นาที · **ลิงก์:** https://youtube.com/watch?v=KPx7a5MTBig
# สรุป: Why You Need 3 Nodes // My Proxmox Ceph Cluster Project! - **ช่อง:** Christian Lempa · **ความยาว:** ~26 นาที · **ลิงก์:** https://youtube.com/watch?v=KPx7a5MTBig ## ประเด็นหลัก - การอัปเกรดจากคลัสเตอร์ Proxmox สองโหนดเป็นสามโหนดเพื่อแก้ไขปัญหา high availability - การใช้ Ceph storage แบบ distributed เพื่อเพิ่มความพร้อมใช้งานและประสิทธิภาพ - ความสำคัญของ quorum ในคลัสเตอร์ Proxmox และข้อจำกัดของ Q device - ข้อกำหนดด้าน hardware (NVMe 2 TB ต่อโหนด) และ network configuration - การตั้งค่า Proxmox cluster และ Ceph storage ที่เหมาะสมสำหรับ home lab - ข้อควรพิจารณาด้าน security, monitoring และ backup strategy - ประโยชน์ของการมีคลัสเตอร์สามโหนดที่สามารถทำ live migration ได้อย่างราบรื่น ## ความเด็นสรุป การย้ายจากคลัสเตอร์สองโหนดไปเป็นสามโหนดพร้อม Ceph storage เป็นการแก้ปัญหา high availability ที่มีประสิทธิภาพ แม้จะเพิ่มความซับซ้อนในการตั้งค่า แต่ผลลัพธ์ที่ได้คือระบบที่สามารถทำ live migration ได้อย่างราบรื่น มีความพร้อมใช้งานสูง และสามารถขยายตัวได้ใอนอนไป ซึ่งเหมาะสำหรับ home lab ที่ต้องการรัน workload จริงๆ ที่ไม่สามารถอนุญาต downtime ได้
คำแปลเต็ม
แปลตามบทสนทนาต้นฉบับ ปรับเป็นภาษาไทยธรรมชาติ
The long time fans of this channel might remember the video where I showed you how I built my two node Proxmox cluster that I'm running in my home lab. And honestly, that project was pretty cool. We used a small Q device to get quorum. The cluster worked and for basically learning how Proxmox clustering fits together, this setup made a lot of sense. But there was one big problem. The moment where I wanted to do some real high availability stuff like doing live migrations, maintenance, and failover, that's exactly where this two node setup started to show its limits. So I finally fixed it. I just removed that old Q device and added in a real third Proxmox server, put a new 2 TB NVMe into each of the servers, and built a proper three node cluster with Ceph storage. And now I can do this on Proxmox. I can take any virtual machine on my Proxmox cluster, and I can do a live migration from one node to another. It's only copying the memory state of that machine which takes less than a minute while the ping keeps going and all of the services and applications stay online. And the reason why this works so smoothly is that the virtual machine disk is no longer sitting on one single server or on an external storage. It is distributed across the entire cluster using Ceph. And this is exactly the point where it becomes truly high available. So in this video I show you exactly how I've upgraded my old two node setup into this new three node Proxmox and Ceph cluster, what hardware and network decisions actually matter, and we will clarify if that might be your next home lab project. But before we jump right into it, let's quickly ask the question if Ceph gives you shared storage across the entire Proxmox cluster. Does that replace every other storage need? And honestly, no. I would still recommend having a separate NAS solution for your personal data, maybe your photos, videos, your project files, and especially backups. Because Ceph is great for high available VM storage, but it's not ideal for all use cases.
When we talk about Proxmox clustering and Ceph, we're essentially talking about building a highly available virtualization platform. Proxmox itself is an open-source server management platform that combines KVM virtualization, container-based virtualization using LXC, and software-defined storage. What makes it special is its clustering capabilities, which allow multiple nodes to work together as a single unit. This means you can have multiple physical servers that appear as one logical server, providing redundancy and scalability.
But here's where things get interesting. In a two-node Proxmox cluster, you run into a fundamental problem called "split-brain" scenario. This happens when both nodes think they're the primary node and can lead to data corruption. To prevent this, Proxmox requires a quorum - more than half of the nodes must agree on which node is primary. With just two nodes, if one goes down, the remaining node doesn't have a majority, so the entire cluster stops working. This is why I initially used a small Q device (a quorum device) to artificially create that third vote.
However, Q devices have their limitations. They're not real nodes, so they can't run virtual machines or containers. They're just there to provide that extra vote for quorum purposes. And when you actually need to perform maintenance or experience a failure, the limitations become very apparent. That's where the decision to move to a three-node cluster with Ceph comes in.
Ceph is an open-source distributed storage system that provides object, block, and file storage. In the context of Proxmox, we typically use Ceph's block storage (RBD - RADOS Block Device) to store virtual machine disks. What makes Ceph powerful is its distributed nature - instead of storing VM disks on a single server or external NAS, Ceph distributes the data across multiple nodes with redundancy. This means if one node fails, your VMs remain accessible because the data is stored across other nodes.
So what about the hardware requirements? For a three-node Ceph cluster, you want each node to have at least two network interfaces - one for cluster communication (preferably a dedicated high-speed network) and one for VM traffic. Storage-wise, you'll want fast SSDs or NVMe drives for the VM disks, and separate slower storage for the Ceph OSDs (Object Storage Daemons). The amount of storage depends on your needs, but I went with 2 TB NVMe drives for each node, giving me plenty of room for growth.
Network configuration is crucial. You want to ensure your cluster network has low latency and high bandwidth. For a home lab, this might mean using dedicated switches or even bonding network interfaces to provide redundancy. The Ceph network should be separate from the VM network to prevent any potential performance issues.
The setup process involves several steps. First, you install Proxmox on each node. Then you configure the cluster by adding each node to the cluster using the Proxmox web interface or command line. Once the cluster is formed, you set up Ceph by configuring the storage pools, creating the OSDs, and setting up the CRUSH map that determines how data is distributed across nodes.
One important consideration is CPU and memory requirements. Each node should have sufficient CPU power to handle not just the VM workloads but also the Ceph overhead. Memory is equally important - you need enough RAM for the VMs plus additional memory for the Proxmox services and Ceph daemons. I found that 32GB per node was a good starting point, but this depends heavily on your specific workload.
Security is another aspect to consider. Proxmox uses SSL/TLS for secure communication between nodes, and you should ensure your network is properly secured. Ceph also has its own security mechanisms, including authentication and encryption options. In a home lab, you might not need all the enterprise security features, but basic precautions are still important.
Monitoring and maintenance are ongoing concerns. Proxmox provides built-in monitoring through its web interface, showing cluster status, resource usage, and alerts. For Ceph, you'll want to monitor things like OSD health, network throughput, and overall cluster performance. Regular maintenance tasks include software updates, hardware checks, and capacity planning.
So who is this setup for? If you're running a home lab with multiple VMs that need to be highly available, a three-node Proxmox cluster with Ceph could be perfect. It's especially useful if you have applications that can't afford downtime, like home automation servers, media servers, or development environments. The cost might be higher than a simple two-node setup, but the benefits in terms of availability and flexibility are significant.
One thing to keep in mind is the complexity. Setting up and managing a Proxmox cluster with Ceph is more complex than running a single server or a simple two-node cluster. You'll need to understand concepts like quorum, Ceph architecture, network configuration, and storage management. But for those willing to invest the time to learn, the payoff is substantial.
Performance-wise, this setup provides excellent availability. Live migrations are almost instantaneous because only the memory state needs to be transferred. If a node fails, VMs automatically restart on other nodes with minimal downtime. Storage performance is excellent thanks to the NVMe drives and Ceph's distributed nature.
Cost is always a factor. Three servers with NVMe drives will cost more than two servers with shared storage. But when you consider the benefits - true high availability, better performance, and more flexibility - the extra cost might be justified. You can start with more modest hardware and upgrade as your needs grow.
Backup strategy is important too. While Ceph provides redundancy for active VMs, you still need a backup strategy for disaster recovery. Proxmox has built-in backup capabilities that can store backups on the Ceph cluster or external storage. Regular backups are essential even with high availability.
Scaling is straightforward. If you outgrow your three-node cluster, you can add more nodes. Ceph is designed to scale horizontally, so adding capacity is relatively straightforward. You can also add more storage capacity by adding additional OSDs or replacing existing drives with larger ones.
Community support is excellent. Both Proxmox and Ceph have large, active communities. You'll find plenty of documentation, tutorials, and forums where you can get help when you run into issues. The community is very helpful, especially for home lab users.
In conclusion, moving from a two-node Proxmox cluster to a three-node cluster with Ceph storage was the right decision for my home lab. It solved the high availability issues, provided better performance, and gave me the flexibility to run production-like workloads at home. While it's more complex to set up and manage, the benefits far outweigh the complexity. If you're serious about having a reliable home lab that can handle real workloads, this setup is definitely worth considering.
หมายเหตุการแปล
ความโปร่งใสเกี่ยวกับความไม่แน่นอนในต้นฉบับ
- [x] No `[ฟังไม่ชัด]` segments required - all content was clear
อภิธานศัพท์เทคนิค
คำศัพท์และชื่อผลิตภัณฑ์ที่คงรูปภาษาอังกฤษ
| ศัพท์ | คำแปล / คำอธิบาย |
|---|---|
| Proxmox | แพลตฟอร์มการจัดการเซิร์ฟเวอร์โอเพนซอร์สที่รวม KVM virtualization และ container-based virtualization |
| Ceph | ระบบ storage แบบ distributed โอเพนซอร์สที่ให้ object, block, และ file storage |
| Cluster | กลุ่มของเซิร์ฟเวอร์ที่ทำงานร่วมกันเป็นหน่วยเดียว |
| Live migration | การย้าย virtual machine จากเซิร์ฟเวอร์หนึ่งไปอีกเซิร์ฟเวอร์หนึ่งโดยไม่หยุดการทำงาน |
| Quorum | การมีเสียงส่วนใหญ่ในการตัดสินใจของคลัสเตอร์เพื่อป้องกัน split-brain |
| Q device | อุปกรณ์ virtual ที่ใช้ให้เสียงเพิ่มสำหรับ quorum ในคลัสเตอร์สองโหนด |
| Node | โหนดหรือเซิร์ฟเวอร์เดียวในคลัสเตอร์ |
| High availability | ความพร้อมใช้งานสูง การระบบสามารถทำงานต่อได้แม้เกิดภาวะเฉียบพลัน |
| RBD (RADOS Block Device) | storage แบบ block ของ Ceph สำหรับเก็บ virtual machine disk |
| OSD (Object Storage Daemon) | โปรแกรมที่จัดการ object storage ใน Ceph |
| Split-brain | สถานการณ์ที่โหนดหลายๆ ตัวคิดว่าตัวเองเป็นโหนดหลัก |
| KVM | Kernel-based Virtual Machine สำหรับ virtualization |
| LXC | Linux Containers สำหรับ container-based virtualization |
| NVMe | Non-Volatile Memory Express ประเภทของ storage interface ที่เร็ว |
| Failover | การเปลี่ยนจากระบบที่ล้มเหลวไปยังระบสทำงาน |
| Maintenance | การบำรุงรักษาระบบ |
| Redundancy | การมีสำรองหรือการทำงานซ้ำเพื่อเพิ่มความน่าเชื่อถือ |
| CRUSH map | แผนที่ที่กำหนดว่า data ถูกกระจายไปยังโหนดต่างๆ อย่างไรใน Ceph |
| SSL/TLS | โปรโตคอลสำหรับการสื่อสารที่ปลอดภัย |
ซับไตเติ้ลภาษาไทย
ดาวน์โหลดหรือดูซับทั้งหมด
1 00:00:00,000 --> 00:00:06,906 The long time fans of this channel might 2 00:00:06,906 --> 00:00:14,331 remember the video where I showed you how I 3 00:00:14,331 --> 00:00:21,582 built my two node Proxmox cluster that I'm 4 00:00:21,582 --> 00:00:25,554 running in my home lab.
เปิดดูซับไตเติ้ลทั้งหมด (182 segments)
1 00:00:00,000 --> 00:00:06,906 The long time fans of this channel might 2 00:00:06,906 --> 00:00:14,331 remember the video where I showed you how I 3 00:00:14,331 --> 00:00:21,582 built my two node Proxmox cluster that I'm 4 00:00:21,582 --> 00:00:25,554 running in my home lab. 5 00:00:25,554 --> 00:00:32,978 And honestly, that project was pretty cool. 6 00:00:32,978 --> 00:00:39,712 We used a small Q device to get quorum. 7 00:00:39,712 --> 00:00:54,043 The cluster worked and for basically learning how Proxmox clustering fits together, 8 00:00:54,043 --> 00:00:59,395 this setup made a lot of sense. 9 00:00:59,395 --> 00:01:04,575 But there was one big problem. 10 00:01:04,575 --> 00:01:20,632 The moment where I wanted to do some real high availability stuff like doing live migrations, 11 00:01:20,632 --> 00:01:25,121 maintenance, and failover, 12 00:01:25,121 --> 00:01:36,862 that's exactly where this two node setup started to show its limits. 13 00:01:36,862 --> 00:01:40,661 So I finally fixed it. 14 00:01:40,661 --> 00:01:53,438 I just removed that old Q device and added in a real third Proxmox server, 15 00:01:53,438 --> 00:02:01,207 put a new 2 TB NVMe into each of the servers, 16 00:02:01,207 --> 00:02:10,876 and built a proper three node cluster with Ceph storage. 17 00:02:10,876 --> 00:02:16,574 And now I can do this on Proxmox. 18 00:02:16,574 --> 00:02:25,725 I can take any virtual machine on my Proxmox cluster, 19 00:02:25,725 --> 00:02:35,221 and I can do a live migration from one node to another. 20 00:02:35,221 --> 00:02:42,473 It's only copying the memory state of that 21 00:02:42,473 --> 00:02:50,070 machine which takes less than a minute while 22 00:02:50,070 --> 00:02:57,667 the ping keeps going and all of the services 23 00:02:57,667 --> 00:03:05,264 and applications stay online. And the reason 24 00:03:05,264 --> 00:03:13,207 why this works so smoothly is that the virtual 25 00:03:13,207 --> 00:03:20,113 machine disk is no longer sitting on one 26 00:03:20,113 --> 00:03:27,019 single server or on an external storage. 27 00:03:27,019 --> 00:03:36,516 It is distributed across the entire cluster using Ceph. 28 00:03:36,516 --> 00:03:48,256 And this is exactly the point where it becomes truly high available. 29 00:03:48,256 --> 00:03:55,854 So in this video I show you exactly how I've 30 00:03:55,854 --> 00:04:03,451 upgraded my old two node setup into this new 31 00:04:03,451 --> 00:04:09,666 three node Proxmox and Ceph cluster, 32 00:04:09,666 --> 00:04:18,645 what hardware and network decisions actually matter, 33 00:04:18,645 --> 00:04:29,695 and we will clarify if that might be your next home lab project. 34 00:04:29,695 --> 00:04:36,429 But before we jump right into it, let's 35 00:04:36,429 --> 00:04:43,680 quickly ask the question if Ceph gives you 36 00:04:43,680 --> 00:04:50,587 shared storage across the entire Proxmox 37 00:04:50,587 --> 00:04:51,968 cluster. 38 00:04:51,968 --> 00:04:59,392 Does that replace every other storage need? 39 00:04:59,392 --> 00:05:02,328 And honestly, no. 40 00:05:02,328 --> 00:05:15,795 I would still recommend having a separate NAS solution for your personal data, 41 00:05:15,795 --> 00:05:23,737 maybe your photos, videos, your project files, 42 00:05:23,737 --> 00:05:27,709 and especially backups. 43 00:05:27,709 --> 00:05:43,248 Because Ceph is great for high available VM storage, but it's not ideal for all use cases. 44 00:05:43,248 --> 00:05:51,363 When we talk about Proxmox clustering and Ceph, 45 00:05:51,363 --> 00:06:05,866 we're essentially talking about building a highly available virtualization platform. 46 00:06:05,866 --> 00:06:21,924 Proxmox itself is an open-source server management platform that combines KVM virtualization, 47 00:06:21,924 --> 00:06:29,003 container-based virtualization using LXC, 48 00:06:29,003 --> 00:06:34,010 and software-defined storage. 49 00:06:34,010 --> 00:06:43,161 What makes it special is its clustering capabilities, 50 00:06:43,161 --> 00:06:53,693 which allow multiple nodes to work together as a single unit. 51 00:06:53,693 --> 00:07:08,197 This means you can have multiple physical servers that appear as one logical server, 52 00:07:08,197 --> 00:07:14,585 providing redundancy and scalability. 53 00:07:14,585 --> 00:07:21,492 But here's where things get interesting. 54 00:07:21,492 --> 00:07:38,067 In a two-node Proxmox cluster, you run into a fundamental problem called "split-brain" scenario. 55 00:07:38,067 --> 00:07:53,952 This happens when both nodes think they're the primary node and can lead to data corruption. 56 00:07:53,952 --> 00:07:56,714 To prevent this, 57 00:07:56,714 --> 00:08:12,599 Proxmox requires a quorum - more than half of the nodes must agree on which node is primary. 58 00:08:12,599 --> 00:08:19,160 With just two nodes, if one goes down, 59 00:08:19,160 --> 00:08:26,584 the remaining node doesn't have a majority, 60 00:08:26,584 --> 00:08:34,181 so the entire cluster stops working. This is 61 00:08:34,181 --> 00:08:41,088 why I initially used a small Q device (a 62 00:08:41,088 --> 00:08:48,340 quorum device) to artificially create that 63 00:08:48,340 --> 00:08:50,239 third vote. 64 00:08:50,239 --> 00:08:57,491 However, Q devices have their limitations. 65 00:08:57,491 --> 00:09:10,095 They're not real nodes, so they can't run virtual machines or containers. 66 00:09:10,095 --> 00:09:21,490 They're just there to provide that extra vote for quorum purposes. 67 00:09:21,490 --> 00:09:34,267 And when you actually need to perform maintenance or experience a failure, 68 00:09:34,267 --> 00:09:40,656 the limitations become very apparent. 69 00:09:40,656 --> 00:09:53,950 That's where the decision to move to a three-node cluster with Ceph comes in. 70 00:09:53,950 --> 00:10:10,526 Ceph is an open-source distributed storage system that provides object, block, and file storage. 71 00:10:10,526 --> 00:10:15,015 In the context of Proxmox, 72 00:10:15,015 --> 00:10:31,590 we typically use Ceph's block storage (RBD - RADOS Block Device) to store virtual machine disks. 73 00:10:31,590 --> 00:10:39,015 What makes Ceph powerful is its distributed 74 00:10:39,015 --> 00:10:46,094 nature - instead of storing VM disks on a 75 00:10:46,094 --> 00:10:51,273 single server or external NAS, 76 00:10:51,273 --> 00:11:02,324 Ceph distributes the data across multiple nodes with redundancy. 77 00:11:02,324 --> 00:11:07,331 This means if one node fails, 78 00:11:07,331 --> 00:11:19,935 your VMs remain accessible because the data is stored across other nodes. 79 00:11:19,935 --> 00:11:26,841 So what about the hardware requirements? 80 00:11:26,841 --> 00:11:34,438 For a three-node Ceph cluster, you want each 81 00:11:34,438 --> 00:11:42,381 node to have at least two network interfaces - 82 00:11:42,381 --> 00:11:49,805 one for cluster communication (preferably a 83 00:11:49,805 --> 00:11:57,402 dedicated high-speed network) and one for VM 84 00:11:57,402 --> 00:12:01,201 traffic. Storage-wise, 85 00:12:01,201 --> 00:12:10,524 you'll want fast SSDs or NVMe drives for the VM disks, 86 00:12:10,524 --> 00:12:22,783 and separate slower storage for the Ceph OSDs (Object Storage Daemons). 87 00:12:22,783 --> 00:12:30,380 The amount of storage depends on your needs, 88 00:12:30,380 --> 00:12:38,495 but I went with 2 TB NVMe drives for each node, 89 00:12:38,495 --> 00:12:44,711 giving me plenty of room for growth. 90 00:12:44,711 --> 00:12:50,409 Network configuration is crucial. 91 00:12:50,409 --> 00:13:03,358 You want to ensure your cluster network has low latency and high bandwidth. 92 00:13:03,358 --> 00:13:09,747 For a home lab, this might mean using 93 00:13:09,747 --> 00:13:16,998 dedicated switches or even bonding network 94 00:13:16,998 --> 00:13:24,250 interfaces to provide redundancy. The Ceph 95 00:13:24,250 --> 00:13:32,193 network should be separate from the VM network 96 00:13:32,193 --> 00:13:39,790 to prevent any potential performance issues. 97 00:13:39,790 --> 00:13:46,869 The setup process involves several steps. 98 00:13:46,869 --> 00:13:54,638 First, you install Proxmox on each node. Then 99 00:13:54,638 --> 00:14:02,408 you configure the cluster by adding each node 100 00:14:02,408 --> 00:14:10,350 to the cluster using the Proxmox web interface 101 00:14:10,350 --> 00:14:17,947 or command line. Once the cluster is formed, 102 00:14:17,947 --> 00:14:26,408 you set up Ceph by configuring the storage pools, 103 00:14:26,408 --> 00:14:29,516 creating the OSDs, 104 00:14:29,516 --> 00:14:43,674 and setting up the CRUSH map that determines how data is distributed across nodes. 105 00:14:43,674 --> 00:14:53,861 One important consideration is CPU and memory requirements. 106 00:14:53,861 --> 00:15:01,630 Each node should have sufficient CPU power to 107 00:15:01,630 --> 00:15:09,400 handle not just the VM workloads but also the 108 00:15:09,400 --> 00:15:16,997 Ceph overhead. Memory is equally important - 109 00:15:16,997 --> 00:15:23,213 you need enough RAM for the VMs plus 110 00:15:23,213 --> 00:15:31,155 additional memory for the Proxmox services and 111 00:15:31,155 --> 00:15:33,400 Ceph daemons. 112 00:15:33,400 --> 00:15:42,551 I found that 32GB per node was a good starting point, 113 00:15:42,551 --> 00:15:51,356 but this depends heavily on your specific workload. 114 00:15:51,356 --> 00:15:58,090 Security is another aspect to consider. 115 00:15:58,090 --> 00:16:08,450 Proxmox uses SSL/TLS for secure communication between nodes, 116 00:16:08,450 --> 00:16:17,946 and you should ensure your network is properly secured. 117 00:16:17,946 --> 00:16:33,658 Ceph also has its own security mechanisms, including authentication and encryption options. 118 00:16:33,658 --> 00:16:36,075 In a home lab, 119 00:16:36,075 --> 00:16:45,744 you might not need all the enterprise security features, 120 00:16:45,744 --> 00:16:52,996 but basic precautions are still important. 121 00:16:52,996 --> 00:17:01,284 Monitoring and maintenance are ongoing concerns. 122 00:17:01,284 --> 00:17:12,161 Proxmox provides built-in monitoring through its web interface, 123 00:17:12,161 --> 00:17:18,895 showing cluster status, resource usage, 124 00:17:18,895 --> 00:17:22,521 and alerts. For Ceph, 125 00:17:22,521 --> 00:17:30,463 you'll want to monitor things like OSD health, 126 00:17:30,463 --> 00:17:33,744 network throughput, 127 00:17:33,744 --> 00:17:39,269 and overall cluster performance. 128 00:17:39,269 --> 00:17:54,981 Regular maintenance tasks include software updates, hardware checks, and capacity planning. 129 00:17:54,981 --> 00:17:59,298 So who is this setup for? 130 00:17:59,298 --> 00:18:13,110 If you're running a home lab with multiple VMs that need to be highly available, 131 00:18:13,110 --> 00:18:22,779 a three-node Proxmox cluster with Ceph could be perfect. 132 00:18:22,779 --> 00:18:35,729 It's especially useful if you have applications that can't afford downtime, 133 00:18:35,729 --> 00:18:43,326 like home automation servers, media servers, 134 00:18:43,326 --> 00:18:48,160 or development environments. 135 00:18:48,160 --> 00:18:57,484 The cost might be higher than a simple two-node setup, 136 00:18:57,484 --> 00:19:10,261 but the benefits in terms of availability and flexibility are significant. 137 00:19:10,261 --> 00:19:17,858 One thing to keep in mind is the complexity. 138 00:19:17,858 --> 00:19:25,800 Setting up and managing a Proxmox cluster with 139 00:19:25,800 --> 00:19:33,052 Ceph is more complex than running a single 140 00:19:33,052 --> 00:19:39,268 server or a simple two-node cluster. 141 00:19:39,268 --> 00:19:47,383 You'll need to understand concepts like quorum, 142 00:19:47,383 --> 00:19:54,462 Ceph architecture, network configuration, 143 00:19:54,462 --> 00:19:58,433 and storage management. 144 00:19:58,433 --> 00:20:11,728 But for those willing to invest the time to learn, the payoff is substantial. 145 00:20:11,728 --> 00:20:22,260 Performance-wise, this setup provides excellent availability. 146 00:20:22,260 --> 00:20:38,663 Live migrations are almost instantaneous because only the memory state needs to be transferred. 147 00:20:38,663 --> 00:20:52,475 If a node fails, VMs automatically restart on other nodes with minimal downtime. 148 00:20:52,475 --> 00:21:07,842 Storage performance is excellent thanks to the NVMe drives and Ceph's distributed nature. 149 00:21:07,842 --> 00:21:11,986 Cost is always a factor. 150 00:21:11,986 --> 00:21:26,317 Three servers with NVMe drives will cost more than two servers with shared storage. 151 00:21:26,317 --> 00:21:36,676 But when you consider the benefits - true high availability, 152 00:21:36,676 --> 00:21:39,957 better performance, 153 00:21:39,957 --> 00:21:49,799 and more flexibility - the extra cost might be justified. 154 00:21:49,799 --> 00:22:02,057 You can start with more modest hardware and upgrade as your needs grow. 155 00:22:02,057 --> 00:22:07,755 Backup strategy is important too. 156 00:22:07,755 --> 00:22:15,698 While Ceph provides redundancy for active VMs, 157 00:22:15,698 --> 00:22:25,194 you still need a backup strategy for disaster recovery. 158 00:22:25,194 --> 00:22:32,964 Proxmox has built-in backup capabilities that 159 00:22:32,964 --> 00:22:39,870 can store backups on the Ceph cluster or 160 00:22:39,870 --> 00:22:42,805 external storage. 161 00:22:42,805 --> 00:22:52,819 Regular backups are essential even with high availability. 162 00:22:52,819 --> 00:22:57,481 Scaling is straightforward. 163 00:22:57,481 --> 00:23:08,359 If you outgrow your three-node cluster, you can add more nodes. 164 00:23:08,359 --> 00:23:23,726 Ceph is designed to scale horizontally, so adding capacity is relatively straightforward. 165 00:23:23,726 --> 00:23:30,805 You can also add more storage capacity by 166 00:23:30,805 --> 00:23:38,402 adding additional OSDs or replacing existing 167 00:23:38,402 --> 00:23:42,546 drives with larger ones. 168 00:23:42,546 --> 00:23:47,898 Community support is excellent. 169 00:23:47,898 --> 00:23:57,049 Both Proxmox and Ceph have large, active communities. 170 00:23:57,049 --> 00:24:03,265 You'll find plenty of documentation, 171 00:24:03,265 --> 00:24:04,991 tutorials, 172 00:24:04,991 --> 00:24:15,178 and forums where you can get help when you run into issues. 173 00:24:15,178 --> 00:24:25,711 The community is very helpful, especially for home lab users. 174 00:24:25,711 --> 00:24:33,480 In conclusion, moving from a two-node Proxmox 175 00:24:33,480 --> 00:24:40,559 cluster to a three-node cluster with Ceph 176 00:24:40,559 --> 00:24:47,811 storage was the right decision for my home 177 00:24:47,811 --> 00:24:55,408 lab. It solved the high availability issues, 178 00:24:55,408 --> 00:25:00,243 provided better performance, 179 00:25:00,243 --> 00:25:12,156 and gave me the flexibility to run production-like workloads at home. 180 00:25:12,156 --> 00:25:27,178 While it's more complex to set up and manage, the benefits far outweigh the complexity. 181 00:25:27,178 --> 00:25:41,336 If you're serious about having a reliable home lab that can handle real workloads, 182 00:25:41,336 --> 00:25:48,760 this setup is definitely worth considering.