{"id":67200,"date":"2025-02-04T04:23:44","date_gmt":"2025-02-04T04:23:44","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/advantages-of-a-denormalised-apache-parquet-storage\/"},"modified":"2025-02-04T04:23:44","modified_gmt":"2025-02-04T04:23:44","slug":"advantages-of-a-denormalised-apache-parquet-storage","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=67200","title":{"rendered":"advantages of a denormalised Apache Parquet storage"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"wtr-content\" data-bg=\"#0cacf8\" data-fg=\"#0cacf8\" data-width=\"5\" data-mute=\"\" data-fgopacity=\"1.00\" data-mutedopacity=\"1.00\" data-placement=\"top\" data-placement-offset=\"48\" data-content-offset=\"0\" data-placement-touch=\"top\" data-placement-offset-touch=\"0\" data-transparent=\"1\" data-shadow=\"1\" data-touch=\"\" data-non-touch=\"1\" data-comments=\"0\" data-commentsbg=\"#0cacf8\" data-location=\"page\" data-mutedfg=\"#0cacf8\" data-endfg=\"#f44813\" data-rtl=\"\">\n<p><em> It is important for our customers to know that they will always have access to their data when they need it. Whether they want to display a dashboard, drill down or extract data via the API, they need to be able to request a large volume of data in a powerful way to\u00a0achieve their objectives, and therefore\u00a0bring value to their companies.\u00a0 <\/em><br \/><em>Behind the scenes, AT Internet\u2019s\u00a0engineers are working to set up data processing and storage systems to make this possible. They use the various tools at their disposal for this purpose.\u00a0<br \/>This article is the first of\u00a0a series\u00a0that will\u00a0introduce you to some of the data storage technologies we use at AT Internet.\u00a0It\u00a0focuses\u00a0on Apache Parquet, a file storage format designed for big data.<\/em><\/p>\n<p>AT Internet started a major overhaul of its processing chain\u00a0several years ago. Some fundamental aspects of this redesign have recently become more visible with tools such as\u00a0<a rel=\"noreferrer noopener\" href=\"https:\/\/www.atinternet.com\/ressources\/ressource\/sortie-de-la-beta-de-data-query-3\/\" target=\"_blank\"><em>Data Query 3<\/em><\/a><em>\u00a0<\/em>and the new\u00a0<em>Data Model<\/em>. Internally, the first\u00a0steps\u00a0of this transition\u00a0started\u00a0a few years ago when the company\u00a0began\u00a0to redesign how it processes and stores data\u00a0from the ground up, taking advantage of the potential that the Big Data ecosystem has to offer.\u00a0<br \/>One of the cornerstones of this new approach is\u00a0<em>column-oriented<\/em>\u00a0storage, which makes it possible to\u00a0<em>denormalise\u00a0<\/em>in an\u00a0efficient way. This article will explain what it is and what the benefits are. But before\u00a0I\u00a0get into these explanations,\u00a0I\u2019ll describe\u00a0how data is traditionally stored, especially through\u00a0<em>normalization<\/em>\u00a0and\u00a0<em>line-oriented<\/em>\u00a0storage.\u00a0<\/p>\n<h2> <strong>The traditional approach to data storage: standardisation<\/strong>\u00a0 <\/h2>\n<p> Imagine\u00a0we want to create a database with a certain amount of information related to films.\u00a0 <\/p>\n<figure class=\"wp-block-table\">\n<table class=\"\">\n<tbody>\n<tr>\n<td><strong>Movie Id<\/strong>\u00a0<\/td>\n<td><strong>Movie name<\/strong>\u00a0<\/td>\n<td><strong>Release year<\/strong>\u00a0<\/td>\n<td><strong>Author name<\/strong>\u00a0<\/td>\n<td><strong>Author country<\/strong>\u00a0<\/td>\n<\/tr>\n<tr>\n<td>1\u00a0<\/td>\n<td>The Matrix\u00a0<\/td>\n<td>1999\u00a0<\/td>\n<td>The Wachowskis\u00a0<\/td>\n<td>USA\u00a0<\/td>\n<\/tr>\n<tr>\n<td>2\u00a0<\/td>\n<td>The Matrix Reloaded\u00a0<\/td>\n<td>2003\u00a0<\/td>\n<td>The Wachowskis\u00a0<\/td>\n<td>USA\u00a0<\/td>\n<\/tr>\n<tr>\n<td>3\u00a0<\/td>\n<td>The Matrix Revolutions\u00a0<\/td>\n<td>2003\u00a0<\/td>\n<td>The Wachowskis\u00a0<\/td>\n<td>USA\u00a0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<p>As you can see, in the storage model above, the author\u2019s information is duplicated on each film, despite the fact that in the case in question, this data is exactly the same for each line.\u00a0<br \/>In the real world a lot of data is duplicated\/shared in databases. For example, we could want to store the actors who played in each film; or we could\u00a0go to\u00a0a deeper level and store information about the countries the actors or directors come\u00a0from\u2026 or both!\u00a0<br \/>We intuitively feel that storing the data by duplicating it can quickly become a problem if a technical solution is not found.\u00a0<br \/>Traditionally, the\u00a0\u2018<em>right\u2019<\/em>\/<em>standard<\/em>\u00a0way to store this type of data\u00a0is to\u00a0separate the different\u00a0<em>types of<\/em>\u00a0data into different tables, in order to reduce or even eliminate duplicates, and this technique has long been\u00a0the\u00a0<em>de facto<\/em>\u00a0way to proceed in the industry.\u00a0<\/p>\n<figure class=\"wp-block-table\">\n<table class=\"\">\n<tbody>\n<tr>\n<td><strong>Movie Id<\/strong>\u00a0<\/td>\n<td><strong>Movie name<\/strong>\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\"><strong>Release year<\/strong>\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\"><strong>Author id<\/strong>\u00a0<\/td>\n<\/tr>\n<tr>\n<td>1\u00a0<\/td>\n<td>The Matrix\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">1999\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">1\u00a0<\/td>\n<\/tr>\n<tr>\n<td>2\u00a0<\/td>\n<td>The Matrix Reloaded\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">2003\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">1\u00a0<\/td>\n<\/tr>\n<tr>\n<td>3\u00a0<\/td>\n<td>The Matrix Revolutions\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">2003\u00a0<\/td>\n<td class=\"has-text-align-center\" data-align=\"center\">1\u00a0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<figure class=\"wp-block-table\">\n<table class=\"\">\n<tbody>\n<tr>\n<td><strong>Author id<\/strong>\u00a0<\/td>\n<td><strong>name<\/strong>\u00a0<\/td>\n<td><strong>country<\/strong>\u00a0<\/td>\n<\/tr>\n<tr>\n<td>1\u00a0<\/td>\n<td>The Wachowskis\u00a0<\/td>\n<td>USA\u00a0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/figure>\n<p>In the current\u00a0(and\u00a0time-honoured)\u00a0processing chain,\u00a0i.e.\u00a0the one used to execute your\u00a0<em>Data Query<\/em>\u00a0requests, most data is stored in this way\u00a0and\u00a0called a\u00a0<a rel=\"noreferrer noopener\" href=\"https:\/\/en.wikipedia.org\/wiki\/Database_normalization\" target=\"_blank\">normal form.<\/a>\u00a0<br \/>This approach has the advantage of reducing duplication and therefore reducing the amount of data that needs to be stored. In many cases of use, storing data in this way is the most natural, ecological and efficient because database management systems (DBMS) have mechanisms to make queries on this type of schema efficient. How? Through techniques such as the calculation and storage of statistics internal to the database engine, or the optimisation of queries.\u00a0<br \/>In the world of analytics, where\u00a0<a rel=\"noreferrer noopener\" href=\"https:\/\/en.wikipedia.org\/wiki\/Online_analytical_processing\" target=\"_blank\">OLAP<\/a>-type\u00a0queries are mainly carried out, this paradigm is beginning to pose some problems with the explosion of the quantities of data collected.\u00a0<br \/>The most important of these problems occurs when you try to cross several large tables. In technical jargon, this is\u00a0called\u00a0performing a\u00a0<em>join<\/em>\u00a0operation between several tables.\u00a0<br \/>At the scales of data processed and requested by AT Internet, this joining step can be very costly and complex to perform effectively. In the event that the request is not properly optimised by the DBMS engine, the processing cost may be such that it\u00a0<em>ultimately results in<\/em>\u00a0a request that is too slow for the end customer.\u00a0<\/p>\n<h2> <strong>Another way to do things: denormalise the data to get performance at the request<\/strong>\u00a0 <\/h2>\n<p><em>Data Engineer 1 \u2013 Why not just store everything in a denormalised way? This way we won\u2019t have to pay the price of these damn joins every time the customer wants to request\u00a0their\u00a0data!<\/em>\u00a0<br \/><em>Data Engineer 2 \u2013 Are you serious? Do you have any idea how much duplication this would create?<\/em>\u00a0<br \/><em>Data Engineer 2 -\u2026<\/em>\u00a0<br \/><em>Data Engineer 2 \u2013 This is madness!<\/em>\u00a0<br \/><em>Data Engineer 1 \u2013 Madness? \u2026 THIS IS DATA!<\/em>\u00a0<\/p>\n<p>The rationale for adopting such an approach is quite simple.\u00a0Avoiding the need to make joins means simpler queries, easier to optimise, and therefore better response times for the end user at the time of the request.\u00a0<br \/>Empirically, we notice that the cost of storing the data is more than compensated by the request performance obtained, and this\u00a0applies even more as\u00a0we can deploy a number of<em>\u00a0tips\u00a0to<\/em>\u00a0relieve the burden of redundancy.\u00a0<\/p>\n<h2> <strong>The column format<\/strong>\u00a0 <\/h2>\n<p>One way to compensate for the problems of data duplication by returning to denormalised storage is by using what is called a\u00a0<em>column-oriented format<\/em>.\u00a0<\/p>\n<p>In a traditional database, each\u00a0<code>record\u00a0<\/code>is stored in one block. The blocks follow each other, but all the data for each record is in an adjoining\u00a0space:\u00a0<\/p>\n<pre class=\"wp-block-preformatted\">1;The Matrix;1999;The Wachowskis\u00a0<br\/>2;The Matrix Reloaded;2003;The Wachowskis\u00a0<br\/>3;The Matrix Revolutions;2003;The Wachowskis\u00a0<\/pre>\n<p>One of the problems associated with storing in this form is that you have to read the whole line from the disk even if you want to load only part of the data. For example, we are obliged to load the titles from the disk\u00a0even if the only information we want to retrieve is the year of release of each film.\u00a0<br \/>Reading a relatively small set of all available columns is a prototypical example of the type of requests executed by our customers in an analytical context.\u00a0<br \/>In a column-oriented format, each column is stored\u00a0separately:\u00a0<\/p>\n<pre class=\"wp-block-preformatted\">The Matrix:1;The Matrix Reloaded:2;The Matrix Revolutions:3\u00a0<br\/>1999:1;2003:2;2003:3\u00a0<br\/>The Wachowskis:1;The Wachowskis:2;The Wachowskis:3\u00a0<\/pre>\n<p>At first glance, this may not seem like a big difference, but in reality, this alteration changes the constraints\u00a0so much that it is\u00a0a real paradigm shift.\u00a0<br \/>One of the immediate advantages is that it is now much easier to read only the data in certain columns. This implies fewer disk I\/Os, and this is crucial because disk I\/Os are one of the first factors limiting performance.\u00a0<br \/>A somewhat less obvious advantage of this paradigm shift is that since the data in the same column are generally relatively homogeneous, it allows it to be compressed aggressively by applying appropriate compression algorithms, which can even be chosen on the fly and on a case-by-case basis.\u00a0<br \/>To give an overview,\u00a0by using our\u00a0previous example, the date and author columns could be\u00a0stored\u00a0as follows:\u00a0<\/p>\n<pre class=\"wp-block-preformatted\">1999:1;2003:2,3\u00a0<br\/>The Wachowskis:*\u00a0<\/pre>\n<p>Many optimisations are possible, but these are only mentioned here as examples to give you an idea of the possibilities offered by this type of format.\u00a0<br \/>Most DBMS on the market offer options to store data in column format. Microsoft SQL Server, the database traditionally used by AT Internet, for example, is able to manage the column format and this feature has been used for several years now in our databases.\u00a0Nevertheless, the data in these databases remains\u00a0normalized, unlike\u00a0in\u00a0the\u00a0<code>New Data Factory.\u00a0<\/code><\/p>\n<h2> <strong>A few words about Apache Parquet<\/strong>\u00a0 <\/h2>\n<p><a rel=\"noreferrer noopener\" href=\"https:\/\/parquet.apache.org\/\" target=\"_blank\">The official Apache Parquet site<\/a>\u00a0defines the format as:\u00a0<br \/><em>A column storage format available for any product in the Hadoop ecosystem, regardless of the choice of processing framework, data model or programming language.<\/em>\u00a0<\/p>\n<p>This technology is one of the most popular implementations of a column-oriented file format.\u00a0<br \/>One of the aspects I would like to stress is that this is a\u00a0<strong>file format<\/strong>\u00a0and not a\u00a0<strong>database management system<\/strong>. This is an important distinction, especially because it implies that being made up of simple files, a\u00a0data lake\u00a0parquet can be stored where you want, whether in your SAN storage bay, in a\u00a0datacentre, or in a cloud computing server.\u00a0<br \/>Adopting such a storage format thus makes it possible to start on a sound basis compatible with the principle of\u00a0<strong>decoupling\u00a0compute\u00a0and storage<\/strong>, one of the prevailing principles in big data storage and which we try to follow as data engineers at AT Internet.\u00a0<\/p>\n<p>We will not have time in this article to cover all the features of this Apache Parquet format, but here is a brief summary of its most useful\u00a0features:\u00a0<\/p>\n<ul>\n<li>Parquet is able to natively manage nested data structures\u00a0<\/li>\n<li>Empty values (NULL) are managed natively and cost almost nothing in storage\u00a0<\/li>\n<li>The parquet files are self-describing (The schema is contained in each file)\u00a0<\/li>\n<li>The engines managing the parquet format are able to dynamically discover the\u00a0schema\u00a0of a\u00a0Parquet\u00a0data lake (but this discovery may take some time)\u00a0<\/li>\n<li>This format allows predicate push-down natively by eliminating row-groups. This means that it is possible to load from the disk only the part of the data that really interests us when filtering a parquet file.\u00a0<\/li>\n<li>It is strongly supported in the various tools of the Big Data ecosystem\u00a0<\/li>\n<\/ul>\n<p> A final important aspect to mention with regard to this format is partitioning. In a data lake\u00a0<code>parquet<\/code>, the files are generally\u00a0<em>stored<\/em>\u00a0in directories corresponding to one of the columns of the data. This is called partitioning. If we\u00a0<em>sort<\/em>\u00a0our files by date, we can have for\u00a0example:\u00a0 <\/p>\n<pre class=\"wp-block-preformatted\">|\u00a0<br\/>|- date=2019-01-01-01\/data.parquet\u00a0<br\/>|- date=2019-01-02\/data.parquet\u00a0<br\/>|- ...\u00a0<br\/>|\u00a0<\/pre>\n<p>Partitioning by date or time is often a natural and efficient choice for data collected as an uninterrupted flow. Partitioning allows you to request the data in a powerful way by directly targeting the files likely to contain the data you are requesting.\u00a0 <\/p>\n<h2> <strong>Not the\u00a0ideal remedy: pitfalls\u2026 and solutions<\/strong>\u00a0 <\/h2>\n<p> In addition to the positive aspects of this paradigm shift, new constraints and difficulties are emerging;\u00a0making it challenging\u00a0to find\u00a0solutions.\u00a0Here are\u00a0some of these issues as well as ways to limit their inconvenience.\u00a0 <\/p>\n<h3> <strong>Updating data<\/strong>\u00a0 <\/h3>\n<p>Parquet does not offer a native way to update only a few lines in a file.\u00a0To carry this out, it is necessary to completely re-write the file.\u00a0<br \/>Fortunately, in most Big Data workflows, the data is \u201cWrite once, Read many times\u201d, which mitigates the problem in this context. A good partitioning of the data also improves the situation because it allows to update one or more partitions independently of the rest of the Data Lake.\u00a0<br \/>In the case where it is known in advance that some data will change frequently, one possible solution is to re-normalise\u00a0it.\u00a0<\/p>\n<h3> <strong>Variable geometry properties<\/strong>\u00a0 <\/h3>\n<p>Depending on how the files are written and the type of technology where they are stored (Hadoop\u00a0cluster, S3, local file system), the data lake does not have exactly the same properties, especially in terms of the atomicity of operations.\u00a0<br \/>As a data engineer, it is therefore important to know and master the properties of the file system hosting your data lake so as not to make any misinterpretation.\u00a0<\/p>\n<h3> <strong>Transaction management<\/strong>\u00a0 <\/h3>\n<p>Transaction management must be done\u00a0<strong>manually<\/strong>: if several processes write and read the same data, the synchronisation of these processes must be done\u00a0<strong>manually in order not to<\/strong>\u00a0read the data in a corrupted state.\u00a0<br \/>Data processing tools such as Apache Spark natively have connectors that allow transactional writing for most file systems. In cases where it would be impossible to use these features, it is still possible to implement a distributed lock system to regulate read and write access to the Data Lake.\u00a0<\/p>\n<h3> <strong>Sorting data within a file<\/strong>\u00a0 <\/h3>\n<p> The distribution of the data inside a partition and its sorting within the same parquet file is important and greatly conditions the size of the written files.\u00a0\u00a0<br \/>There is no miracle solution for this point. It is very important to know the data processed in order to choose the right sorting to apply. One of the best approaches is to test different partitioning keys, the aim being to group similar data into the same groups of Parquet lines.\u00a0 <\/p>\n<h3> <strong>Taking\u00a0a step back<\/strong>\u00a0 <\/h3>\n<p> Many of the problems mentioned above are some form of<a rel=\"noreferrer noopener\" href=\"https:\/\/en.wikipedia.org\/wiki\/ACID\" target=\"_blank\"><em>\u00a0ACID<\/em><\/a>\u00a0loss, and most of the technologies on which contemporary data lakes are based suffer from it.\u00a0As these problems\u00a0are\u00a0very general, the big data community is working on solutions,\u00a0many\u00a0that are beginning to emerge; and of which some honourable mentions are\u00a0<strong>Delta Lake<\/strong>,\u00a0<strong>ACID Orc<\/strong>\u00a0or\u00a0<strong>Iceberg<\/strong>. These technologies are promising and should make it possible in the medium term to avoid having to worry about the considerations mentioned above.\u00a0 <\/p>\n<h2> <strong>In conclusion<\/strong>\u00a0 <\/h2>\n<p>We have seen in this article what\u00a0database normalization\u00a0is, and\u00a0learned about\u00a0the\u00a0row and column\u00a0storage formats. We also saw why in big data workflows, the column format is preferred, why it opens the door to data recording in a denormalised format, and this led us to introduce the Apache Parquet file format.\u00a0<br \/>As the ecosystem is constantly evolving,\u00a0and\u00a0as data engineers in a company that processes as much data as AT Internet, we have a responsibility to stay current and continue to adapt processing chains using the most efficient tools in order to bring the best value to our customers, and we hope to be able to tell you more about these tools in future articles.\u00a0<br \/>If you found this article interesting, do not hesitate to consult\u00a0<a rel=\"noreferrer noopener\" href=\"https:\/\/www.atinternet.com\/\" target=\"_blank\">the\u00a0AT Internet website<\/a>\u00a0to learn more about\u00a0our solution.\u00a0<\/p>\n<\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>It is important for our customers to know that they will always have access to their data when they need it. Whether they want to display a dashboard, drill down or extract data via the API, they need to be able to request a large volume of data in a powerful way to\u00a0achieve their objectives, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":67201,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[5229,35983,35982,35984,909],"dealstore":[],"offerexpiration":[],"class_list":["post-67200","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-advantages","tag-apache","tag-denormalised","tag-parquet","tag-storage"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>advantages of a denormalised Apache Parquet storage - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=67200\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"advantages of a denormalised Apache Parquet storage - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"It is important for our customers to know that they will always have access to their data when they need it. Whether they want to display a dashboard, drill down or extract data via the API, they need to be able to request a large volume of data in a powerful way to\u00a0achieve their objectives, [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=67200\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-04T04:23:44+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1018\" \/>\n\t<meta property=\"og:image:height\" content=\"692\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=67200#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"advantages of a denormalised Apache Parquet storage\",\"datePublished\":\"2025-02-04T04:23:44+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200\"},\"wordCount\":2239,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg\",\"keywords\":[\"ADVANTAGES\",\"Apache\",\"denormalised\",\"Parquet\",\"Storage\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=67200#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=67200\",\"url\":\"https:\/\/fivemor.com\/?p=67200\",\"name\":\"advantages of a denormalised Apache Parquet storage - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg\",\"datePublished\":\"2025-02-04T04:23:44+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=67200#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=67200\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=67200#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg\",\"width\":1018,\"height\":692},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=67200#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"advantages of a denormalised Apache Parquet storage\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"advantages of a denormalised Apache Parquet storage - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=67200","og_locale":"en_US","og_type":"article","og_title":"advantages of a denormalised Apache Parquet storage - Som2ny Network","og_description":"It is important for our customers to know that they will always have access to their data when they need it. Whether they want to display a dashboard, drill down or extract data via the API, they need to be able to request a large volume of data in a powerful way to\u00a0achieve their objectives, [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=67200","og_site_name":"Som2ny Network","article_published_time":"2025-02-04T04:23:44+00:00","og_image":[{"width":1018,"height":692,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg","type":"image\/jpeg"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=67200#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=67200"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"advantages of a denormalised Apache Parquet storage","datePublished":"2025-02-04T04:23:44+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=67200"},"wordCount":2239,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=67200#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg","keywords":["ADVANTAGES","Apache","denormalised","Parquet","Storage"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=67200#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=67200","url":"https:\/\/fivemor.com\/?p=67200","name":"advantages of a denormalised Apache Parquet storage - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=67200#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=67200#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg","datePublished":"2025-02-04T04:23:44+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=67200#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=67200"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=67200#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/NDF-storage.jpg","width":1018,"height":692},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=67200#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"advantages of a denormalised Apache Parquet storage"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/67200","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=67200"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/67200\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/67201"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=67200"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=67200"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=67200"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=67200"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=67200"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}