{
    "componentChunkName": "component---node-modules-lekoarts-gatsby-theme-minimal-blog-core-src-templates-post-query-tsx",
    "path": "/why-is-blue-store-the-way-it-is-part-1-why-file-systems-fall-short",
    "result": {"data":{"post":{"slug":"/why-is-blue-store-the-way-it-is-part-1-why-file-systems-fall-short","title":"Why Is BlueStore the Way It Is? Part 1: Why File Systems Fall Short","date":"09.10.2026","tags":[{"name":"ceph","slug":"ceph"},{"name":"bluestore","slug":"bluestore"},{"name":"storage","slug":"storage"},{"name":"distributed-systems","slug":"distributed-systems"}],"description":null,"canonicalUrl":null,"body":"var _excluded = [\"components\"];\nfunction _extends() { return _extends = Object.assign ? Object.assign.bind() : function (n) { for (var e = 1; e < arguments.length; e++) { var t = arguments[e]; for (var r in t) ({}).hasOwnProperty.call(t, r) && (n[r] = t[r]); } return n; }, _extends.apply(null, arguments); }\nfunction _objectWithoutProperties(e, t) { if (null == e) return {}; var o, r, i = _objectWithoutPropertiesLoose(e, t); if (Object.getOwnPropertySymbols) { var s = Object.getOwnPropertySymbols(e); for (r = 0; r < s.length; r++) o = s[r], t.includes(o) || {}.propertyIsEnumerable.call(e, o) && (i[o] = e[o]); } return i; }\nfunction _objectWithoutPropertiesLoose(r, e) { if (null == r) return {}; var t = {}; for (var n in r) if ({}.hasOwnProperty.call(r, n)) { if (e.includes(n)) continue; t[n] = r[n]; } return t; }\n/* @jsxRuntime classic */\n/* @jsx mdx */\n\nvar _frontmatter = {\n  \"title\": \"Why Is BlueStore the Way It Is? Part 1: Why File Systems Fall Short\",\n  \"date\": \"2026-10-09T00:00:00.000Z\",\n  \"draft\": false,\n  \"tags\": [\"ceph\", \"bluestore\", \"storage\", \"distributed-systems\"]\n};\nvar layoutProps = {\n  _frontmatter: _frontmatter\n};\nvar MDXLayout = \"wrapper\";\nreturn function MDXContent(_ref) {\n  var components = _ref.components,\n    props = _objectWithoutProperties(_ref, _excluded);\n  return mdx(MDXLayout, _extends({}, layoutProps, props, {\n    components: components,\n    mdxType: \"MDXLayout\"\n  }), mdx(\"p\", null, \"Hey folks, as a continuation of my goal to understand BlueStore more deeply, I\\nam writing my second blog post, revolving around the question: \\\"Why is BlueStore\\nthe way it is?\\\" When reading the BlueStore code and resources, the following\\nquestions kept bugging me:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Why use a bespoke solution like BlueStore? Why are we not using already\\nexisting solutions like file systems?\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"What were the previous solutions that were considered before we zeroed in on\\nthe current architecture?\")), mdx(\"p\", null, \"I personally understand concepts better when I get some historical context that\\nled to the present. The quote \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"\\\"Those who cannot remember the past are condemned\\nto repeat it\\\"\"), \" forms one of the bases of my learning structure. This is what led\\nme on the quest for the origins of BlueStore. Fortunately, I stumbled across the\\npaper titled\\n\", mdx(\"a\", {\n    parentName: \"p\",\n    \"href\": \"https://dl.acm.org/doi/10.1145/3341301.3359656\"\n  }, \"\\\"File Systems Unfit as Distributed Storage Backends: Lessons from 10 Years of Ceph Evolution\\\"\"), \"\\n(Aghayev et al., SOSP '19).\"), mdx(\"p\", null, \"I encourage everyone to read the paper. This post is hugely motivated by the\\npaper and thus contains some excerpts from it.\"), mdx(\"p\", null, \"In this series we will explore the evolution of the storage backends that were\\nused in Ceph, how the problems with each of them motivated the next version, and\\nhow all of the learnings combined paved the way for BlueStore. This turned out\\nto be a long one, so I have split it into three parts:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Part 1 (this post):\"), \" What a distributed system needs from its storage\\nbackend, and why file systems struggle to provide it.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Part 2:\"), \" Ceph's backends in order (EBOFS \\u2192 FileStore on Btrfs \\u2192 FileStore\\non XFS \\u2192 NewStore) and how each of them ran into these problems.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Part 3:\"), \" How all of these learnings shaped BlueStore, and the price it\\npays for owning the disk.\")), mdx(\"h2\", null, \"Intro\"), mdx(\"p\", null, \"A storage backend is defined as the software module directly managing the\\nstorage device attached to a physical machine, i.e., it's the software on each\\nmachine that actually puts the bytes on the disk. These backends are entirely\\nlocal. In a distributed system, where there are many physical storage devices,\\nit is the role of the system to aggregate the storage backends on the many\\nphysical devices and present them as a single unified data store. Note that the\\ndistributed layer's guarantees are built on top of the backend's guarantees.\"), mdx(\"p\", null, \"Side note: when the paper says \\\"Distributed Storage Backends\\\", read it as\\n(Distributed Storage) Backends, i.e., the backends that are used in the\\ndistributed storage system.\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \" \\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500 Ceph / RADOS (distributed system) \\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n \\u2502       CRUSH, PGs, replication, recovery, cluster maps     \\u2502\\n \\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u252C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u252C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u252C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n         \\u2502                 \\u2502                 \\u2502\\n  OSD.0 on node A   OSD.1 on node B   OSD.2 on node C\\n         \\u2502                 \\u2502                 \\u2502\\n   BlueStore #0      BlueStore #1      BlueStore #2    \\u2190 \\\"distributed storage backends\\\"\\n         \\u2502                 \\u2502                 \\u2502           (three independent, local instances)\\n     /dev/sdb          /dev/sdb          /dev/sdb\\n\")), mdx(\"h2\", null, \"What Does a Distributed System Need from Its Backend?\"), mdx(\"p\", null, \"In the terminology of Ceph, RADOS can only promise strong consistency if each\\nOSD can apply a multi-op transaction atomically.\"), mdx(\"p\", null, \"Why? Because a single client write is a transaction that updates several things\\nat once:\"), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"The object's data\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"The object's metadata\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"An entry in the placement group's log\")), mdx(\"p\", null, \"When a crash happens, RADOS checks which replicas are up to date by comparing\\nthe logs. This makes sure that the client doesn't receive stale data. This is\\nonly possible if each OSD's log tells the truth about its data. If a crash could\\nleave the log entry for version 42 on disk while the data is still at version\\n41, the OSD would claim that it has the current data while still holding stale\\nbytes, and RADOS would believe it, serving the client old data.\"), mdx(\"p\", null, \"Hence the backend must apply each transaction as all-or-nothing on its own disk.\\nKeeping replicas in agreement is RADOS's job. Each backend must just make sure\\nthat what it reports about itself is true.\"), mdx(\"p\", null, \"It is my understanding that a good backend has the following properties:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Strong consistency\"), \", which needs atomicity, durability, ordering and\\nintegrity.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Efficiency\"), \", which needs cheap transactions and fast metadata.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Longevity\"), \", which needs hardware flexibility.\")), mdx(\"p\", null, \"Each of these properties is hard to get on top of a file system, and this maps\\nnicely to the three problems we will look at later in this post:\"), mdx(\"table\", null, mdx(\"thead\", {\n    parentName: \"table\"\n  }, mdx(\"tr\", {\n    parentName: \"thead\"\n  }, mdx(\"th\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"A good backend needs\"), mdx(\"th\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Which is hard on a file system because of\"))), mdx(\"tbody\", {\n    parentName: \"table\"\n  }, mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Strong consistency\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, mdx(\"strong\", {\n    parentName: \"td\"\n  }, \"Transactions\"), \": there are no multi-op atomic transactions\")), mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Efficiency\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, mdx(\"strong\", {\n    parentName: \"td\"\n  }, \"Transactions\"), \" (costly to fake) and \", mdx(\"strong\", {\n    parentName: \"td\"\n  }, \"Metadata\"), \" (slow with millions of objects)\")), mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Longevity\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, mdx(\"strong\", {\n    parentName: \"td\"\n  }, \"New hardware\"), \": we have to wait for the file system to support it\")))), mdx(\"h2\", null, \"File Systems as Storage Backends\"), mdx(\"p\", null, \"Historically, storage began with people/companies owning the disk. As one might\\nimagine, it caused a lot of issues. Later, when local file systems were\\ndeveloped, storage adopted them, as they were mature, familiar and often good\\nenough for the large-file workloads of the 2000s. With changing workloads (small\\nobjects and transactions) and advancements in hardware (NVMe, zones), we are\\noutgrowing what a general-purpose file system can offer cheaply and are\\nreturning to owning the disks.\"), mdx(\"p\", null, \"Ceph's own path is the whole story in miniature:\"), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"It started with EBOFS, a custom user-space object store\\n(\", mdx(\"a\", {\n    parentName: \"li\",\n    \"href\": \"https://github.com/ceph/ceph/commit/18d9132a053846d7fd24ec512a83d4da106d6ed7\"\n  }, \"added in 2005\"), \").\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"It moved to FileStore on Btrfs (~2008) to get transactions, checksums and\\ncompression.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Then XFS, for stability.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Then back to the raw disk with BlueStore.\")), mdx(\"p\", null, \"We'll explore each of them in detail in Part 2.\"), mdx(\"p\", null, \"File systems became the de facto standard for the following reasons:\"), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Delegate the hard parts.\"), \" Data persistence, crash-safe metadata, allocation\\nand caching come from well-tested, mature and highly performant code. A team\\nbuilding a distributed system need not reinvent the wheel.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"A familiar interface.\"), \" Files, directories and POSIX are easy to reason\\nabout.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Tooling.\"), \" Standard tools such as \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"ls\"), \" and \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"find\"), \" can be used to explore the\\ndisk contents.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Two environmental reasons:\", mdx(\"ul\", {\n    parentName: \"li\"\n  }, mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Linux's ubiquity.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"The workload fit: the workloads of that time did not require millions of\\nsmall objects, overwrites or multi-op atomicity.\")))), mdx(\"h2\", null, \"Problems with File Systems as Storage Backends\"), mdx(\"p\", null, \"For about 10 years, after a brief start with EBOFS, Ceph built its storage\\nbackend on local file systems (FileStore on Btrfs \\u2192 FileStore on XFS \\u2192\\nNewStore).\"), mdx(\"p\", null, \"There are three main problems with using file systems as the storage backend for\\ndistributed systems:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Inefficient transactions\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Slow metadata operations\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Less support for new storage hardware\")), mdx(\"p\", null, \"We will explore each of these issues in this post, and in Part 2 we will map\\nwhich backend had which issues and how they motivated the design of BlueStore.\"), mdx(\"h3\", null, \"Problem 1: Transactions\"), mdx(\"p\", null, \"As we saw in the earlier section, a distributed system's consistency guarantees\\nare built on top of the guarantees of its storage backends. This means each\\nbackend must be able to apply multi-op transactions atomically, so that after a\\ncrash the local state is the same as the one it reports to the cluster. Hence,\\natomicity (all-or-nothing) during a transaction is a hard requirement.\"), mdx(\"p\", null, \"POSIX file systems have no transaction APIs. This means there is no way to\\nexpress something like \\\"apply these three operations atomically\\\". They give us\\nindividual atomic operations (\", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"rename\"), \", \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"open\"), \", etc.) but no multi-op\\ntransactions, and durability comes only from \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"fsync\"), \". File systems do use\\ntransactions internally to keep their own metadata consistent, but generally\\ndon't expose them to applications.\"), mdx(\"p\", null, \"For example, in order to atomically replace a file, we need to do something like\\nthis:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\",\n    \"className\": \"language-c\"\n  }, \"// 1. Write the new version to a temporary file\\nint fd = open(\\\"config.tmp\\\", O_WRONLY | O_CREAT | O_TRUNC, 0644);\\nwrite(fd, new_contents, len);\\n\\n// 2. Make the new contents durable\\nfsync(fd);\\nclose(fd);\\n\\n// 3. Atomically swap it into place\\nrename(\\\"config.tmp\\\", \\\"config\\\");\\n\\n// 4. Make the swap itself durable\\nint dirfd = open(\\\".\\\", O_RDONLY);\\nfsync(dirfd);\\nclose(dirfd);\\n\")), mdx(\"p\", null, \"Notice that after each step we call \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"fsync\"), \" to make it durable. This is fine if\\nthere is only one thing to switch, i.e., the directory entry that maps the name\\n\", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"config\"), \" to an inode. A storage backend like Ceph's needs a single write to\\nupdate the object's data, its metadata and a log entry, all together.\"), mdx(\"p\", null, \"One cannot do the following:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"op1: write object data (v42 bytes)\\nfsync\\nop2: set xattr version = 42\\nfsync\\nop3: append PG log entry 42\\nfsync\\n\")), mdx(\"p\", null, \"If a crash happens at any of these steps, the intermediate data is wrong, and\\nit's hard to roll back.\"), mdx(\"p\", null, \"Hence a backend that is built on a file system has to construct transactions on\\ntop of it. There are three ways to do it, and Ceph tried all of them:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Hooking into the file system's internal transaction mechanism (FileStore on\\nBtrfs)\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Implementing a write-ahead log (WAL) in user space (FileStore on XFS)\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Using a key-value database with transactions as a WAL (NewStore)\")), mdx(\"p\", null, \"The next sections talk about why these options have significant\\nperformance/complexity overhead.\"), mdx(\"h4\", null, \"Leveraging the File System's Internal Transactions\"), mdx(\"p\", null, \"Many file systems have an in-kernel transaction framework. This is a mechanism\\nthe file system uses so that its own multi-step updates are crash-safe. For\\nexample, when you create a file, the file system updates the free-space bitmap,\\nthe inode and the directory entry, and wraps them in an internal transaction so\\nthat a crash can't leave them half-done.\"), mdx(\"p\", null, \"A sample in-kernel transaction (from ext4's journal, jbd2) looks like:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\",\n    \"className\": \"language-c\"\n  }, \"handle = jbd2_journal_start(journal, nblocks);  // reserve journal space\\n    // modify bitmap block in memory\\n    // modify inode block in memory\\n    // modify directory block in memory\\njbd2_journal_stop(handle);\\n\")), mdx(\"p\", null, \"The idea is to expose this machinery to user space so that applications like\\nCeph can get atomic multi-op transactions for free.\"), mdx(\"p\", null, mdx(\"strong\", {\n    parentName: \"p\"\n  }, \"Why these internal transactions aren't enough\")), mdx(\"p\", null, \"These transactions were built for a different job, i.e., to keep the file\\nsystem's own structure consistent. Hence they have the following properties:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Not exposed.\"), \" Applications can't normally use them at all.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"No rollback.\"), \" Kernel file operations are designed to never fail halfway.\\nBefore modifying anything, the kernel does all the checks that could fail\\n(permissions, free space, journal credits, etc.) to prevent mid-operation\\nfailures.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"No clear boundaries.\"), \" Some file systems batch many operations from many\\nprocesses into one big transaction every few seconds. There is no notion of\\nsegregation, i.e., nothing that says \\\"these three ops from this process, and\\nonly these\\\".\")), mdx(\"p\", null, \"The property that bites Ceph is \\\"no rollback\\\". Let's say we expose these\\ninternal transactions for applications to use. We could then do something like:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"TRANS_START\\n  op1 \\u2705 applied\\n                  \\u2190 process dies\\n  op2 never issued\\n  op3 never issued\\nTRANS_END never issued\\n\")), mdx(\"p\", null, \"Since this is an in-kernel transaction started from user space, the kernel now\\nholds an open transaction whose owner is gone. It cannot undo op1, because it\\nhas no rollback mechanism. Its only option is to close the transaction as-is and\\ncommit op1 alone.\"), mdx(\"p\", null, \"This leads to a state where the object on disk has v42 data but a v41 version\\nand a v41 log entry, which is an inconsistent state.\"), mdx(\"p\", null, \"But why can't the kernel just undo op1? This is because file system journals\\nlike jbd2 are redo-only, i.e., they only record the \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"new\"), \" contents of each block\\nand never the old ones. So there is nothing to undo with. Rollback was never\\nneeded in the first place, since kernel operations validate everything up front\\nand don't fail halfway. If something truly unexpected happens in the middle of a\\ntransaction, the file system does not roll back; it aborts the journal and goes\\nread-only (e.g., ext4's \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"errors=remount-ro\"), \", a Btrfs transaction abort).\"), mdx(\"p\", null, \"Note that this only bites when the \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"process\"), \" dies and the kernel survives. If\\nthe whole machine crashes, the uncommitted kernel transaction simply vanishes,\\nwhich is fine. The dangerous case is when the OSD process dies (or hits\\n\", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"ENOSPC\"), \") while the kernel keeps running and commits the half-done transaction.\"), mdx(\"h4\", null, \"Implementing the WAL in User Space\"), mdx(\"p\", null, \"A logical write-ahead log in user space is an alternative to using the file\\nsystem's in-kernel transaction framework. It provides atomicity for transactions\\nby writing down the full plan before doing any real work. Once the whole plan is\\nsafely on the disk, it adds one tiny \\\"commit\\\" record at the end. The disk writes\\nthis commit record completely or not at all.\"), mdx(\"p\", null, \"If the system crashes, recovery just checks for the commit record. If it's\\nthere, recovery finishes the job. If it isn't, recovery ignores the plan, and\\nsince no real work had started, nothing needs undoing. Either way, we never get\\nstuck with a half-done change.\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \" 1. Write the plan      2. Write COMMIT     3. Do the work\\n[ A-500, B+500 ] \\u2500\\u2500\\u2500> [ COMMIT \\u2714 ] \\u2500\\u2500\\u2500> [ change A and B ]\\n\\nCrash? Check the log:\\n  COMMIT missing \\u2500\\u2500> ignore plan \\u2500\\u2500> nothing happened\\n  COMMIT present \\u2500\\u2500> redo plan   \\u2500\\u2500> everything happened\\n\")), mdx(\"p\", null, \"In terms of Ceph, the storage backend maintains its own WAL, called the journal.\\nWhen a write comes in, the following happens:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"The transaction is serialized and written to the journal.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"fsync\"), \" is called to commit the transaction to disk.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"The operations in the transaction are applied to the disk.\")), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"Transaction t = {\\n    write(obj A, 4 KiB @ 8192),\\n    setattr(A, v42),\\n    append(pglog, entry 42)\\n}\\n\\nStep 1: Serialize t (including the 4 KiB of data) \\u2192 append to journal\\nStep 2: fsync the journal   \\u2190 COMMIT POINT. Now t is durable (safe).\\nStep 3: Apply t to XFS: pwrite(A), fsetxattr(A), pwrite(pglog)\\n\")), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Crash before step 2:\"), \" the journal entry is incomplete, so ignore it. Treat\\nit as if the transaction never happened.\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Crash after step 2:\"), \" on restart, we replay all the journal entries and they\\nget reapplied. Say the crash happens at \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"fsetxattr(A)\"), \": when the replay\\nhappens during recovery, we begin the entire transaction again from\\n\", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"pwrite(A)\"), \".\")), mdx(\"p\", null, \"The journal gives us atomicity (one commit point) and durability. But it has\\nthree costs.\"), mdx(\"p\", null, mdx(\"strong\", {\n    parentName: \"p\"\n  }, \"Cost 1: Slow read-modify-write\")), mdx(\"p\", null, \"Ceph workloads usually follow a read-modify-write pattern. This is because the\\nclient often writes a smaller unit of data, but redundancy, checksums, etc. are\\nall maintained over something bigger: a stripe, a blob, an omap entry set or an\\nobject version. Ceph has to bring that bigger unit up to date, and doing that\\nrequires knowing its current state. That is a read, then a modify, then a write.\"), mdx(\"p\", null, \"The issue with the WAL is that a transaction's effect cannot be read by a\\nsubsequent transaction until it has been committed to the journal (the \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"fsync\"), \"\\nstep) \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"and\"), \" applied to the file system. This is because FileStore serves reads\\nfrom the file system, not from the journal.\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"txn 1: increment counter in object X (10 \\u2192 11)\\ntxn 2: increment counter in object X (needs to read 11)\\n\\ntime \\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u25BA\\ntxn 1: [serialize][journal write][fsync ~ms][apply to XFS]\\ntxn 2:                                                    [read X = 11]\\n                                                          [serialize][journal][fsync][apply]\\n\")), mdx(\"p\", null, \"Every dependent operation now has to wait for the entire commit and apply before\\nit can proceed.\"), mdx(\"p\", null, mdx(\"strong\", {\n    parentName: \"p\"\n  }, \"Cost 2: Non-idempotent operations\")), mdx(\"p\", null, \"Most operations done on the ObjectStore are idempotent; they are just set\\noperations (e.g., \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"write(X, offset, data)\"), \", \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"truncate\"), \", \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"setattr\"), \"). There are a\\nfew non-idempotent operations, such as \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"clone(a \\u2192 b)\"), \", \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"clone_range\"), \", \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"remove\"), \"\\nand \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"split_collection\"), \". Each of these reads live state (data) and writes based\\non it.\"), mdx(\"p\", null, \"After a crash, replaying operations is only safe if each produces the same\\nresult when applied twice, but that's not the case with non-idempotent\\noperations, since they depend on the live state of the object.\"), mdx(\"p\", null, \"The flow of transactions is like this:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Write the transaction into the journal.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Apply it to the file system.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Some time later, the journal is trimmed and the old entries are thrown away.\")), mdx(\"p\", null, \"At any moment, recent transactions are in one of three states:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"txn 900\\u2013950    applied + synced      \\u2192 trimmable, gone after next sync\\ntxn 951\\u2013990    applied, NOT synced   \\u2192 still in the journal (the risky window)\\ntxn 991\\u20131000   journaled, not applied yet\\n\")), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"applied\"), \" = data has been written into the page cache (RAM)\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"synced\"), \" = \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"syncfs\"), \", the flush mechanism that moves data from the page cache\\nonto the physical disk\")), mdx(\"p\", null, \"When a transaction is \\\"applied\\\", the data usually lands in the Linux page cache\\nin RAM, not on disk. The file system writes it out whenever it chooses. So\\n\\\"applied\\\" doesn't mean \\\"durable\\\".\"), mdx(\"p\", null, \"If the machine crashes, replay starts from the last sync point, i.e., txn 951.\\nTransactions 951\\u2013990 might already be partly or fully on disk, and they might\\nget applied again. That's where non-idempotent operations cause issues.\"), mdx(\"p\", null, \"For example:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"txn:\\n \\u2460 clone a\\u2192b\\n \\u2461 update a\\n \\u2462 update c\\n\")), mdx(\"p\", null, \"Say a crash occurs after \\u2461 has reached the disk. The physical storage now holds\\nthe updated value of \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"a\"), \". When replay starts again from \\u2460, the clone copies the\\nupdated value of \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"a\"), \" into \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"b\"), \" instead of the original value.\"), mdx(\"p\", null, mdx(\"strong\", {\n    parentName: \"p\"\n  }, \"Cost 3: Double writes\")), mdx(\"p\", null, \"The journal is a full write-ahead log. The whole transaction, including the data\\npayload, is serialized and written to the journal and then applied to the file\\nsystem. Every byte of data is written twice:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"once into the journal\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"later into the file system\")), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"4 MiB write \\u2192 4 MiB into journal (fsync) \\u2192 4 MiB into XFS\\n\")), mdx(\"h4\", null, \"Using a Key-Value Store as the WAL\"), mdx(\"p\", null, \"In this approach, we put the transaction state into a key-value store rather\\nthan writing and managing our own log. This way, we hand off the hard part (the\\nWAL machinery) to libraries like RocksDB, and the system just reads and writes\\nkeys.\"), mdx(\"p\", null, \"In a hand-rolled WAL we need to design all of this:\"), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"a log file format\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"a method to serialize transactions into it\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"commit records + \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"fsync\")), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"replay logic after a crash (this must be idempotent)\")), mdx(\"p\", null, \"When we use a KV store as the WAL, we can just write code like this:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"batch = new WriteBatch();\\nbatch.put(\\\"obj/A/info\\\",   {version: 42, size: 12288});\\nbatch.put(\\\"obj/A/attr/_\\\", ...);\\nbatch.put(\\\"pglog/42\\\",     \\\"modify A v42\\\");\\ndb.write(batch, sync=true);   // atomic + durable, done\\n\")), mdx(\"p\", null, \"The KV store does exactly what we would have built. From the user's point of\\nview, all we need to know is that once the write returns, the batch is durable\\nand applied all-or-nothing.\"), mdx(\"p\", null, \"The following are the conceptual changes:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"We now store state, not a log of operations.\"), \" In a WAL, we recorded\\noperations that were applied elsewhere; in a KV store, we write the resulting\\nstate to the associated key.\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"The KV store becomes the source of truth.\"), \" Since we have keys, it's easy\\nto answer questions like \\\"Does object A exist?\\\", \\\"Where's its data?\\\", \\\"What\\nversion is it?\\\" You can just look up the key instead of inferring the answer\\nfrom file names and file sizes.\")), mdx(\"p\", null, \"Example: a tiny object store\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"Keys in RocksDB:\\n  obj/photo1 \\u2192 {data_at: extent 500\\u2013520, size: 81920, version: 3}\\n  obj/photo2 \\u2192 {data_at: extent 900\\u2013902, size: 12000, version: 1}\\n\")), mdx(\"p\", null, \"Writing a new photo:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"1. allocate extents 1200\\u20131240, write the data there, flush\\n2. db.write({ put obj/photo3 \\u2192 {data_at: 1200\\u20131240, ...} }, sync)\\n\")), mdx(\"p\", null, \"A crash before step 2 leaves extents 1200\\u20131240 unreferenced: free space that a\\ncleanup pass reclaims. A crash after step 2 means the photo fully exists.\"), mdx(\"p\", null, mdx(\"strong\", {\n    parentName: \"p\"\n  }, \"Cost: Journal on top of a journal\")), mdx(\"p\", null, \"The KV store keeps a diary (log) so that it can recover after a crash. But this\\nKV store runs on top of a file system, and the file system also keeps its own\\ndiary to protect its files. So every time the KV store says \\\"save my diary\\nentry\\\", two diaries get updated. These disk flushes are slow, often milliseconds\\non a hard drive, and having two flushes cuts how many objects we can create per\\nsecond.\"), mdx(\"p\", null, \"When the object's data also lives in a file on the same file system, creating\\none object looks like this:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"fsync(object's data file)\\n  \\u251C\\u2500 flush the object's data     \\u2190 flush 1\\n  \\u2514\\u2500 flush XFS's own journal     \\u2190 flush 2 (records the new file and its size)\\nfsync(RocksDB's log file)\\n  \\u251C\\u2500 flush RocksDB's log entry   \\u2190 flush 3\\n  \\u2514\\u2500 flush XFS's own journal     \\u2190 flush 4 (records \\\"this log file grew\\\")\\n\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\ntotal: 4 flushes (on a raw disk: 2)\\n\")), mdx(\"h4\", null, \"Summary of the three approaches\"), mdx(\"p\", null, \"To summarize the three ways of building transactions on top of a file system:\"), mdx(\"table\", null, mdx(\"thead\", {\n    parentName: \"table\"\n  }, mdx(\"tr\", {\n    parentName: \"thead\"\n  }, mdx(\"th\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Approach\"), mdx(\"th\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Atomicity comes from\"), mdx(\"th\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"What it costs\"))), mdx(\"tbody\", {\n    parentName: \"table\"\n  }, mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Borrow the kernel's transactions\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"The file system's internal journal\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"No rollback: a dead process leaves half a transaction committed\")), mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"User-space WAL\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Our own journal + \", mdx(\"inlineCode\", {\n    parentName: \"td\"\n  }, \"fsync\")), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Slow read-modify-write, unsafe replay, every byte written twice\")), mdx(\"tr\", {\n    parentName: \"tbody\"\n  }, mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"KV store as WAL\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"A RocksDB write batch\"), mdx(\"td\", {\n    parentName: \"tr\",\n    \"align\": null\n  }, \"Journal on a journal: 4 flushes where a raw disk needs 2\")))), mdx(\"h3\", null, \"Problem 2: Metadata Operations\"), mdx(\"p\", null, \"Metadata is information about the data that is stored. It contains everything we\\nneed to know about a piece of data in order to find it, manage it and trust it.\"), mdx(\"p\", null, \"For a single 4 MB RADOS object, we have the following metadata:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Object name\"), \", e.g. \", mdx(\"inlineCode\", {\n    parentName: \"li\"\n  }, \"rbd_data.1f2a.00032\"), \", and the PG and collection it\\nbelongs to\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Attributes\"), \": size, version, snapshot, info, xattrs, omap keys\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Location\"), \": the disk extents that hold its bytes, which OSD\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, mdx(\"strong\", {\n    parentName: \"li\"\n  }, \"Integrity and history\"), \": checksums, PG log entry\")), mdx(\"p\", null, \"This metadata may run to a few hundred bytes. Because metadata is what keeps the\\ndata findable and trustworthy, every data operation also updates it; a single\\ndata operation might require 4 or more metadata updates.\"), mdx(\"p\", null, \"In distributed storage systems, where a lot of operations are performed on data,\\ninefficient metadata operations become a bottleneck.\"), mdx(\"p\", null, \"The objects on the file system would be stored as shown below:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"/var/lib/ceph/osd/ceph-0/current/\\n\\u2514\\u2500\\u2500 2.3f_head/                                   \\u2190 one directory = one PG\\n    \\u251C\\u2500\\u2500 rbd\\\\udata.abc...09__head_0C11B03F__2\\n    \\u251C\\u2500\\u2500 rbd\\\\udata.abc...05__head_E5326A3F__2\\n    \\u251C\\u2500\\u2500 rbd\\\\udata.def...01__head_7A40913F__2\\n    \\u251C\\u2500\\u2500 rbd\\\\udata.ghi...02__head_51C0D23F__2\\n    \\u251C\\u2500\\u2500 ...\\n    \\u2514\\u2500\\u2500 ... 1,000,000 files in the same directory\\n\")), mdx(\"p\", null, \"And the problem with storing metadata in the file system can be captured by the\\ndiagram below:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"/var/lib/ceph/osd/ceph-0/current/2.3f_head/\\n\\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n\\u2502 rbd\\\\udata.abc...09__head_0C11B03F__2 \\u2192 inode 8812 \\u2500\\u2500\\u2510        \\u2502\\n\\u2502 rbd\\\\udata.abc...05__head_E5326A3F__2 \\u2192 inode 1201   \\u2502        \\u2502\\n\\u2502 rbd\\\\udata.def...01__head_7A40913F__2 \\u2192 inode 4410   \\u2502        \\u2502\\n\\u2502 rbd\\\\udata.ghi...02__head_51C0D23F__2 \\u2192 inode 7733   \\u2502        \\u2502\\n\\u2502 ... 1,000,000 more entries ...                      \\u2502        \\u2502\\n\\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u253C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n        \\u2502                                             \\u25BC\\n        \\u2502                         inode 8812 (metadata)\\n        \\u2502                         \\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n        \\u2502                         \\u2502 size: 4 MB                   \\u2502\\n        \\u2502                         \\u2502 owner, permissions, times    \\u2502\\n        \\u2502                         \\u2502 xattrs: version, attrs,      \\u2502\\n        \\u2502                         \\u2502   (real name if too long)    \\u2502\\n        \\u2502                         \\u2502 data: blocks 9000\\u201310023 \\u2500\\u2500\\u2500\\u2500\\u2500\\u253C\\u2500\\u2500\\u25BA object bytes\\n        \\u2502                         \\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n        \\u2502\\n        \\u2502 \\u2460 readdir(): returns ALL 1,000,000 entries,\\n        \\u2502    in the file system's own order (random to Ceph)\\n        \\u25BC\\n\\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n\\u2502 ..._0C11B03F, ..._E5326A3F, ..._7A40913F \\u2502 \\u2190 unsorted\\n\\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n        \\u2502\\n        \\u2502 \\u2461 long names? read each inode's xattrs\\n        \\u2502    to get the real object name (1 inode lookup per entry)\\n        \\u25BC\\n\\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n\\u2502 sort 1,000,000 names by hash             \\u2502\\n\\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n        \\u2502\\n        \\u25BC\\n\\u250C\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2510\\n\\u2502 objects in hash order                    \\u2502 \\u2190 what Ceph actually needed\\n\\u2514\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2500\\u2518\\n\")), mdx(\"p\", null, \"Enumerating objects is a very important operation for a distributed system, both\\nfor keeping data consistent and for repairs.\"), mdx(\"p\", null, \"For Ceph, enumeration means \\\"listing all objects that exist in a placement\\ngroup\\\". Ceph has no central metadata server holding the list of all objects.\\nEach OSD only knows what it holds. So whenever the system needs answers to:\"), mdx(\"ul\", null, mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Which objects is this replica missing?\"), mdx(\"li\", {\n    parentName: \"ul\"\n  }, \"Do all replicas hold the same objects with the same contents?\")), mdx(\"p\", null, \"the only way to answer is to ask each OSD to list what it has and compare the\\nlists.\"), mdx(\"p\", null, \"Notice that one directory per PG already gives us only that PG's objects. The\\nreal problem is the \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"order\"), \". Scrub and backfill don't ask for the whole list at\\nonce; they walk it in chunks, by hash:\"), mdx(\"pre\", null, mdx(\"code\", {\n    parentName: \"pre\"\n  }, \"scrub: \\\"give me the next 100 objects after hash 0x3A000000\\\"\\n       \\u2192 compare with the replica's next 100 \\u2192 repeat\\n\")), mdx(\"p\", null, \"Hash order lets two replicas walk their lists in lockstep and compare them chunk\\nby chunk. It also lets a scrub pause and resume, locking only the range it is\\ncurrently checking. \", mdx(\"inlineCode\", {\n    parentName: \"p\"\n  }, \"readdir()\"), \" can't do this, since it returns the entries in\\nthe file system's own order. So answering \\\"the next 100 objects after H\\\" means\\nreading and sorting the \", mdx(\"em\", {\n    parentName: \"p\"\n  }, \"entire\"), \" directory every single time, as shown in the\\ndiagram above.\"), mdx(\"h3\", null, \"Problem 3: New Storage Hardware\"), mdx(\"p\", null, \"Newer drives like SMR HDDs and ZNS SSDs don't allow random overwrites; the\\nsoftware must write sequentially in zones. A backend built on a file system\\ncan't use them until the file system learns to, and that can take years. I am\\nskipping this one in depth for this blog, since I need to understand it better\\nmyself.\"), mdx(\"h2\", null, \"Wrapping Up\"), mdx(\"p\", null, \"A distributed system can only be as correct as its backends. Ceph needs each\\nOSD's backend to do three things well:\"), mdx(\"ol\", null, mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Apply multi-op transactions atomically and cheaply\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Handle metadata fast, including listing millions of objects in order\"), mdx(\"li\", {\n    parentName: \"ol\"\n  }, \"Adopt new storage hardware without waiting on someone else\")), mdx(\"p\", null, \"File systems don't give us any of these cheaply. We have to fake transactions on\\ntop of them, the metadata ends up spread across directories, inodes and xattrs\\nthat can't be listed in the order Ceph needs, and new hardware has to wait until\\nthe file system supports it.\"), mdx(\"p\", null, \"In Part 2, we will see Ceph run into exactly these problems, one backend at a\\ntime, from EBOFS to NewStore.\"));\n}\n;\nMDXContent.isMDXComponent = true;","excerpt":"Hey folks, as a continuation of my goal to understand BlueStore more deeply, I\nam writing my second blog post, revolving around the question…","timeToRead":12,"banner":null}},"pageContext":{"slug":"/why-is-blue-store-the-way-it-is-part-1-why-file-systems-fall-short","formatString":"DD.MM.YYYY"}},
    "staticQueryHashes": ["2744905544","3090400250","318001574"]}