{"id":101701,"date":"2025-02-21T15:20:34","date_gmt":"2025-02-21T15:20:34","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/scrape-the-urls-of-a-domain-and-write-the-results-to-bigquery\/"},"modified":"2025-02-21T15:20:34","modified_gmt":"2025-02-21T15:20:34","slug":"scrape-the-urls-of-a-domain-and-write-the-results-to-bigquery","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=101701","title":{"rendered":"Scrape The URLs Of A Domain And Write The Results To BigQuery"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>In my intense love affair with the <a href=\"https:\/\/cloud.google.com\/\">Google Cloud Platform<\/a>, I\u2019ve never felt more inspired to write content and try things out. After starting with a <a href=\"https:\/\/www.simoahava.com\/analytics\/install-snowplow-on-the-google-cloud-platform\/\">Snowplow Analytics setup guide<\/a>, and continuing with a <a href=\"https:\/\/www.simoahava.com\/google-cloud\/lighthouse-bigquery-google-cloud-platform\/\">Lighthouse audit automation tutorial<\/a>, I\u2019m going to show you yet another cool thing you can do with GCP.<\/p>\n<div style=\"aspect-ratio: 981 \/ 501;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/bigquery-status-codes.jpg\" title=\"BigQuery status codes\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"501\" width=\"981\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/bigquery-status-codes.jpg#ZgotmplZ\" alt=\"BigQuery status codes\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>In this guide, I\u2019ll show you how to use an <a href=\"https:\/\/github.com\/yujiosaka\/headless-chrome-crawler\">open-source web crawler<\/a> running in a <a href=\"https:\/\/cloud.google.com\/compute\/\">Google Compute Engine<\/a> virtual machine (VM) instance to scrape all the internal and external links of a given domain, and write the results into a <a href=\"https:\/\/cloud.google.com\/bigquery\/\">BigQuery table<\/a>. With this setup, you can audit and monitor the links in any website, looking for bad status codes or missing titles, and fix them to improve your site\u2019s logical architecture.<\/p>\n<p>                <span class=\"simmer\"><br \/>\n  <span class=\"close\">X<\/span><\/p>\n<p>\n    <span class=\"fa fa-md fa-bell\"\/><br \/>\n    <strong>The Simmer Newsletter<\/strong>\n  <\/p>\n<p>\n    Subscribe to the <a href=\"https:\/\/www.simoahava.com\/newsletter\/\">Simmer newsletter<\/a> to get the latest news and content from Simo Ahava into your email inbox!\n  <\/p>\n<p>  <\/span><\/p>\n<h2 id=\"how-it-works\">How it works<\/h2>\n<p>The idea is fairly simple. You\u2019re using a <strong>Google Compute Engine<\/strong> VM instance to run the crawler script. The purpose here is that you can scale the instance up as much as you like (and can afford) to get the extra power you might not have with your local machine.<\/p>\n<div style=\"aspect-ratio: 1020 \/ 266;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/compute-engine-instance.jpg\" title=\"Compute engine instance\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"266\" width=\"1020\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/compute-engine-instance.jpg#ZgotmplZ\" alt=\"Compute engine instance\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>The crawler runs through the pages of the domain you specify in the <strong>configuration<\/strong>, and writes the results into a <strong>BigQuery<\/strong> table.<\/p>\n<p>There are only a few moving parts here. Whenever you want to run the crawl again, all you need to do is just start the instance again. You won\u2019t be charged for the time the instance is stopped (the script auto-stops the instance once the crawl is done), so you can simply leave the instance in its stopped state until you need to do a recrawl.<\/p>\n<p>You could even create a <a href=\"https:\/\/cloud.google.com\/functions\/\">Google Cloud Function<\/a> that starts the instance with a trigger (an HTTP request or a <a href=\"https:\/\/cloud.google.com\/pubsub\/docs\/\">Pub\/Sub message<\/a>, for example). There are many ways to skin this cat, too!<\/p>\n<p>The configuration also has a setting for utilizing a <a href=\"https:\/\/redis.io\/\">Redis<\/a> cache by way of <a href=\"https:\/\/cloud.google.com\/memorystore\/\">GCP Memorystore<\/a>, for example. The cache is useful if you have a huuuuuuge domain to crawl and you want to be able to pause\/resume the crawl, or even utilize more than one VM instance to do the crawl.<\/p>\n<p>The <strong>cost<\/strong> of running this setup really depends on how big the crawl is and how much power you dedicate to the VM instance.<\/p>\n<p>On my own site, the ~7500 links and images being crawled take about 10 minutes on a 16 CPU, 60 GB instance (without Redis) VM instance. This translates to around 50 cents per crawl. I could scale down the instance for a lower cost, and I\u2019m sure there are other ways of optimizing it, too.<\/p>\n<h2 id=\"preparations\">Preparations<\/h2>\n<p>The preparations are almost the same as in my <a href=\"https:\/\/www.simoahava.com\/google-cloud\/lighthouse-bigquery-google-cloud-platform\/\">earlier articles<\/a>, but with some simplifications.<\/p>\n<h3 id=\"install-command-line-tools\">Install command line tools<\/h3>\n<p>Start by installing the following CLI tools:<\/p>\n<ol>\n<li>\n<p><a href=\"https:\/\/cloud.google.com\/sdk\/\">Google Cloud SDK<\/a><\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/git-scm.com\/downloads\">Git<\/a><\/p>\n<\/li>\n<\/ol>\n<p>To verify you have these up and running, run the following commands in your terminal:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-shell\" data-lang=\"shell\">$ gcloud -v\nGoogle Cloud SDK 228.0.0\n\n$ git --version\ngit version 2.19.2<\/code><\/pre>\n<\/div>\n<h3 id=\"set-up-a-new-google-cloud-platform-project-with-billing\">Set up a new Google Cloud Platform project with Billing<\/h3>\n<p>Follow the steps <a href=\"https:\/\/www.simoahava.com\/google-cloud\/lighthouse-bigquery-google-cloud-platform\/#set-up-a-new-google-cloud-platform-project-with-billing\">here<\/a>, and make sure you write down the Project ID since you\u2019ll need it in a number of places. I\u2019ll use my example of <code>web-scraper-gcp<\/code> in this guide.<\/p>\n<div style=\"aspect-ratio: 811 \/ 361;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/web-scraper-gcp-project.jpg\" title=\"New project\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"361\" width=\"811\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/web-scraper-gcp-project.jpg#ZgotmplZ\" alt=\"New project\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"clone-the-github-repo-and-edit-the-configuration\">Clone the Github repo and edit the configuration<\/h3>\n<p>Before getting things up and running in GCP, you\u2019ll need to create a <strong>configuration<\/strong> file first.<\/p>\n<p>The easiest way to access the necessary files is to clone the <a href=\"https:\/\/www.github.com\/\">Github<\/a> repo for this project.<\/p>\n<ol>\n<li>\n<p>Browse to a local directory where you want to write the contents of the repo to.<\/p>\n<\/li>\n<li>\n<p>Run the following command to write the files into a new folder named <code>web-scraper-gcp\/<\/code>:<\/p>\n<\/li>\n<\/ol>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-shell\" data-lang=\"shell\">$ git clone https:\/\/github.com\/sahava\/web-scraper-gcp.git<\/code><\/pre>\n<\/div>\n<div style=\"aspect-ratio: 977 \/ 248;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/web-scraper-dir.jpg\" title=\"Web scraper directory\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"248\" width=\"977\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/web-scraper-dir.jpg#ZgotmplZ\" alt=\"Web scraper directory\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Next, run the command <code>mv config.json.sample config.json<\/code> while in the <code>web-scraper-gcp\/<\/code> directory.<\/p>\n<p>Finally, open the file <code>config.json<\/code> for editing in your favorite text editor. Here\u2019s what the sample file looks like:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-json\" data-lang=\"json\">{\n  <span style=\"color:#1e90ff;font-weight:bold\">\"domain\"<\/span>: <span style=\"color:#a50\">\"www.gtmtools.com\"<\/span>,\n  <span style=\"color:#1e90ff;font-weight:bold\">\"startUrl\"<\/span>: <span style=\"color:#a50\">\"https:\/\/www.gtmtools.com\/\"<\/span>,\n  <span style=\"color:#1e90ff;font-weight:bold\">\"projectId\"<\/span>: <span style=\"color:#a50\">\"web-scraper-gcp\"<\/span>,\n  <span style=\"color:#1e90ff;font-weight:bold\">\"bigQuery\"<\/span>: {\n    <span style=\"color:#1e90ff;font-weight:bold\">\"datasetId\"<\/span>: <span style=\"color:#a50\">\"web_scraper_gcp\"<\/span>,\n    <span style=\"color:#1e90ff;font-weight:bold\">\"tableId\"<\/span>: <span style=\"color:#a50\">\"crawl_results\"<\/span>\n  },\n  <span style=\"color:#1e90ff;font-weight:bold\">\"redis\"<\/span>: {\n    <span style=\"color:#1e90ff;font-weight:bold\">\"active\"<\/span>: <span style=\"color:#00a\">false<\/span>,\n    <span style=\"color:#1e90ff;font-weight:bold\">\"host\"<\/span>: <span style=\"color:#a50\">\"10.0.0.3\"<\/span>,\n    <span style=\"color:#1e90ff;font-weight:bold\">\"port\"<\/span>: <span style=\"color:#099\">6379<\/span>\n  }\n}<\/code><\/pre>\n<\/div>\n<p>Here\u2019s an explanation of what the fields are and what you\u2019ll need to do.<\/p>\n<table>\n<thead>\n<tr>\n<th>Field<\/th>\n<th>Value<\/th>\n<th>Description<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>\"domain\"<\/code><\/td>\n<td><code>\"gtmtools.com\"<\/code><\/td>\n<td>This is used for determining what is an <strong>internal<\/strong> and what is an <strong>external<\/strong> URL. The check will be a pattern match, so if the crawled URL <em>includes<\/em> this string, it will be considered an internal URL.<\/td>\n<\/tr>\n<tr>\n<td><code>\"startUrl\"<\/code><\/td>\n<td><code>\"https:\/\/www.gtmtools.com\/\"<\/code><\/td>\n<td>A fully qualified URL address that represents the entry point of the crawl.<\/td>\n<\/tr>\n<tr>\n<td><code>\"projectId\"<\/code><\/td>\n<td><code>\"web-scraper-gcp\"<\/code><\/td>\n<td>The Google Cloud Platform project ID.<\/td>\n<\/tr>\n<tr>\n<td><code>\"bigQuery.datasetId\"<\/code><\/td>\n<td><code>\"web_scraper_gcp\"<\/code><\/td>\n<td>The ID of the BigQuery dataset the script will attempt to create. You must follow the <a href=\"https:\/\/cloud.google.com\/bigquery\/docs\/datasets#create-dataset\">naming rules<\/a>.<\/td>\n<\/tr>\n<tr>\n<td><code>\"bigQuery.tableId\"<\/code><\/td>\n<td><code>\"crawl_results\"<\/code><\/td>\n<td>The ID of the table the script will attempt to create. You must follow the <a href=\"https:\/\/cloud.google.com\/bigquery\/docs\/tables#create-table\">naming rules<\/a>.<\/td>\n<\/tr>\n<tr>\n<td><code>\"redis.active\"<\/code><\/td>\n<td><code>false<\/code><\/td>\n<td>Set to <code>true<\/code> if you want to use a <strong>Redis<\/strong> cache to persist the crawl queue.<\/td>\n<\/tr>\n<tr>\n<td><code>\"redis.host\"<\/code><\/td>\n<td><code>\"10.0.0.3\"<\/code><\/td>\n<td>Set to the IP address via which the script can connect to the Redis instance.<\/td>\n<\/tr>\n<tr>\n<td><code>\"redis.port\"<\/code><\/td>\n<td><code>6379<\/code><\/td>\n<td>Set to the port number of the Redis instance (usually <code>6379<\/code>).<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Once you\u2019ve edited the configuration, you\u2019ll need to upload it to a Google Cloud Storage bucket.<\/p>\n<h3 id=\"upload-the-configuration-to-gcs\">Upload the configuration to GCS<\/h3>\n<p>Browse to <a href=\"https:\/\/console.cloud.google.com\/storage\/browser\">https:\/\/console.cloud.google.com\/storage\/browser<\/a> and make sure you have the correct project selected.<\/p>\n<div style=\"aspect-ratio: 899 \/ 268;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/gcs-correct-project.jpg\" title=\"GCS correct project\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"268\" width=\"899\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/gcs-correct-project.jpg#ZgotmplZ\" alt=\"GCS correct project\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Next, <strong>create a new bucket<\/strong> in a region nearby, and give it an easy-to-remember name.<\/p>\n<div style=\"aspect-ratio: 913 \/ 597;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/new-gcs-bucket-nearline.jpg\" title=\"New GCS bucket\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"597\" width=\"913\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/new-gcs-bucket-nearline.jpg#ZgotmplZ\" alt=\"New GCS bucket\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Once done, enter the bucket, choose <strong>Upload files<\/strong>, and locate the <code>config.json<\/code> file from your local computer and upload it into the bucket.<\/p>\n<div style=\"aspect-ratio: 991 \/ 375;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/config-json-in-bucket.jpg\" title=\"Upload file to GCS\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"375\" width=\"991\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/config-json-in-bucket.jpg#ZgotmplZ\" alt=\"Upload file to GCS\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"edit-the-install-script\">Edit the install script<\/h3>\n<p>The Git repo you downloaded comes with a file named <code>gce-install.sh<\/code>. This script will be used to fire up the VM instance with the correct settings (and it will initiate the crawl when started). However, you\u2019ll need to edit the file so that it knows where to fetch your configuration file from. So, <strong>open<\/strong> the <code>gce-install.sh<\/code> file for editing.<\/p>\n<p>Edit the following line:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-shell\" data-lang=\"shell\"><span style=\"color:#a00\">bucket<\/span>=<span style=\"color:#a50\">'gs:\/\/web-scraper-config\/config.json'<\/span><\/code><\/pre>\n<\/div>\n<p>Change the <code>web-scraper-config<\/code> part to the name of the bucket you just created. So if you named the bucket <code>my-configuration-bucket<\/code>, you\u2019d change the line to this:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-shell\" data-lang=\"shell\"><span style=\"color:#a00\">bucket<\/span>=<span style=\"color:#a50\">'gs:\/\/my-configuration-bucket\/config.json'<\/span><\/code><\/pre>\n<\/div>\n<h3 id=\"make-sure-the-required-services-have-been-enabled-in-google-cloud-platform\">Make sure the required services have been enabled in Google Cloud Platform<\/h3>\n<p>Final preparatory step is to make sure you have the required services enabled in Google Cloud Platform.<\/p>\n<ol>\n<li>\n<p>Browse <a href=\"https:\/\/console.cloud.google.com\/apis\/api\/compute.googleapis.com\">here<\/a>, and make sure the Compute Engine API has been enabled.<\/p>\n<\/li>\n<li>\n<p>Browse <a href=\"https:\/\/console.cloud.google.com\/apis\/api\/bigquery-json.googleapis.com\">here<\/a>, and make sure the BigQuery API has been enabled.<\/p>\n<\/li>\n<li>\n<p>Browse <a href=\"https:\/\/console.cloud.google.com\/apis\/api\/storage-component.googleapis.com\">here<\/a>, and make sure the Google Cloud Storage API has been enabled.<\/p>\n<\/li>\n<\/ol>\n<h2 id=\"create-the-gce-vm-instance\">Create the GCE VM instance<\/h2>\n<p>Now you\u2019re ready to create the Google Compute Engine instance, firing it up with your install script. Here\u2019s what\u2019s going to happen when you do so:<\/p>\n<ol>\n<li>\n<p>Once the instance is created, it will run the <code>gce-install.sh<\/code> script. In fact, it will run this script whenever you start the instance again.<\/p>\n<\/li>\n<li>\n<p>The script will install all the required dependencies to run the web crawler. There\u2019s quite a few of them because running a headless Chrome browser in a virtual machine isn\u2019t the most trivial operation.<\/p>\n<\/li>\n<li>\n<p>The penultimate step of the install script is to run the Node app containing the code I\u2019ve written to perform the crawl task.<\/p>\n<\/li>\n<li>\n<p>The Node app will grab the <code>startUrl<\/code> and BigQuery information from the configuration file (downloaded from the GCS bucket), and will crawl the domain, writing the results into BigQuery.<\/p>\n<\/li>\n<li>\n<p>Once the crawl is complete, the VM instance will shut itself down.<\/p>\n<\/li>\n<\/ol>\n<p>To create the instance, you\u2019ll need to run this command:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-shell\" data-lang=\"shell\">$ gcloud compute instances create web-scraper-gcp <span style=\"color:#a50\">\\\n<\/span><span style=\"color:#a50\"\/>      --metadata-from-file=startup-script=.\/gce-install.sh <span style=\"color:#a50\">\\\n<\/span><span style=\"color:#a50\"\/>      --scopes=bigquery,cloud-platform <span style=\"color:#a50\">\\\n<\/span><span style=\"color:#a50\"\/>      --machine-type=n1-standard-16 <span style=\"color:#a50\">\\\n<\/span><span style=\"color:#a50\"\/>      --zone=europe-north1-a<\/code><\/pre>\n<\/div>\n<p>Edit the <code>machine-type<\/code> and <code>zone<\/code> if you want to have the instance run on a different CPU\/memory profile, and\/or if you want to run it in a different zone. You can find a list of the machine types <a href=\"https:\/\/cloud.google.com\/compute\/docs\/machine-types\">here<\/a>, and a list of zones <a href=\"https:\/\/cloud.google.com\/compute\/docs\/regions-zones\/\">here<\/a>.<\/p>\n<p>Once done, you should see something like this:<\/p>\n<div style=\"aspect-ratio: 1302 \/ 72;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/gce-create-command.jpg\" title=\"GCE Create command\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"72\" width=\"1302\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/gce-create-command.jpg#ZgotmplZ\" alt=\"GCE Create command\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h2 id=\"check-to-see-if-it-works\">Check to see if it works<\/h2>\n<p>First, head on over to the <a href=\"https:\/\/console.cloud.google.com\/compute\/instances\">instance list<\/a>, and make sure you see the instance running (you should see a green checkmark next to it):<\/p>\n<div style=\"aspect-ratio: 1197 \/ 256;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/vm-instance-running.jpg\" title=\"VM instance running\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"256\" width=\"1197\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/vm-instance-running.jpg#ZgotmplZ\" alt=\"VM instance running\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Naturally, the fact that it\u2019s running doesn\u2019t really tell you much, yet.<\/p>\n<p>Next, head on over to <a href=\"https:\/\/console.cloud.google.com\/bigquery\">BigQuery<\/a>. You should see your project in the navigator, so click it open. Under the project, you should see a dataset and a table.<\/p>\n<div style=\"aspect-ratio: 887 \/ 297;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/dataset-table.jpg\" title=\"Dataset and table\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"297\" width=\"887\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/dataset-table.jpg#ZgotmplZ\" alt=\"Dataset and table\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>If you see them, the next step is to run a simple query in the <strong>query editor<\/strong>. Click the table name in the navigator, and then click the <strong>QUERY TABLE<\/strong> link. The Query editor should be pre-filled with a table query, so between the <code>SELECT<\/code> and <code>FROM<\/code> keywords, type: <code>count(*)<\/code>. This is what the query should end up looking like:<\/p>\n<div style=\"aspect-ratio: 1504 \/ 517;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/select-count.jpg\" title=\"Select count(*)\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"517\" width=\"1504\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/select-count.jpg#ZgotmplZ\" alt=\"Select count(*)\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Finally, click the <strong>Run<\/strong> button.<\/p>\n<p>This will run the query against the BigQuery table. The crawl is still probably running, but thanks to <strong>streaming inserts<\/strong> it is constantly adding rows to the table. The query should return a result that shows you how many rows there are in the table currently:<\/p>\n<div style=\"aspect-ratio: 628 \/ 277;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/test-query.jpg\" title=\"Test the query\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"277\" width=\"628\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/test-query.jpg#ZgotmplZ\" alt=\"Test the query\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>If you see a result, it means the whole thing is working! Keep on monitoring the size of the table. Once the crawl finishes, the virtual machine instance will shut down, and you\u2019ll be able to see it in its stopped state.<\/p>\n<div style=\"aspect-ratio: 1025 \/ 259;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/12\/stopped-instance.jpg\" title=\"Stopped instance\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"259\" width=\"1025\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/12\/stopped-instance.jpg#ZgotmplZ\" alt=\"Stopped instance\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h2 id=\"final-thoughts\">Final thoughts<\/h2>\n<p>First of all, this was an <strong>exercise<\/strong>. I\u2019m fully aware of awesome crawling tools such as <a href=\"https:\/\/www.screamingfrog.co.uk\/seo-spider\/\">Screaming Frog<\/a>, which you can use to achieve very much the same thing.<\/p>\n<p>However, this setup has some cool features:<\/p>\n<ol>\n<li>\n<p>You can modify the crawler with additional <a href=\"https:\/\/github.com\/yujiosaka\/headless-chrome-crawler\/blob\/master\/docs\/API.md\">options<\/a>, AND you can pass <a href=\"https:\/\/github.com\/GoogleChrome\/puppeteer\/blob\/v1.11.0\/docs\/api.md\">flags<\/a> to the <a href=\"https:\/\/github.com\/GoogleChrome\/puppeteer\">Puppeteer<\/a> instance running in the background.<\/p>\n<\/li>\n<li>\n<p>Since this crawler uses a headless browser, it works better on dynamically generated sites than a regular HTTP request crawler. It actually generates the JavaScript and crawls the dynamic links, too.<\/p>\n<\/li>\n<li>\n<p>Because it writes the data to BigQuery, you can monitor the status codes and link integrity of your website in tools like Google Data Studio.<\/p>\n<\/li>\n<\/ol>\n<p>Anyway, I didn\u2019t set out to create a tool that replaces some of the stuff already out there. Instead, I wanted to show you how easy it is to run scripts and perform tasks in the Google Cloud.<\/p>\n<p>Let me know in the comments if you\u2019re having trouble with this setup! I\u2019m happy to see where the problem might lie.<\/p>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>In my intense love affair with the Google Cloud Platform, I\u2019ve never felt more inspired to write content and try things out. After starting with a Snowplow Analytics setup guide, and continuing with a Lighthouse audit automation tutorial, I\u2019m going to show you yet another cool thing you can do with GCP. In this guide, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":101702,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[28608,11109,4955,47142,14426,4012],"dealstore":[],"offerexpiration":[],"class_list":["post-101701","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-bigquery","tag-domain","tag-results","tag-scrape","tag-urls","tag-write"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=101701\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"In my intense love affair with the Google Cloud Platform, I\u2019ve never felt more inspired to write content and try things out. After starting with a Snowplow Analytics setup guide, and continuing with a Lighthouse audit automation tutorial, I\u2019m going to show you yet another cool thing you can do with GCP. In this guide, [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=101701\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-21T15:20:34+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"981\" \/>\n\t<meta property=\"og:image:height\" content=\"501\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"9 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=101701#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Scrape The URLs Of A Domain And Write The Results To BigQuery\",\"datePublished\":\"2025-02-21T15:20:34+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701\"},\"wordCount\":1666,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg\",\"keywords\":[\"BigQuery\",\"domain\",\"Results\",\"Scrape\",\"URLs\",\"Write\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=101701#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=101701\",\"url\":\"https:\/\/fivemor.com\/?p=101701\",\"name\":\"Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg\",\"datePublished\":\"2025-02-21T15:20:34+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=101701#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=101701\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=101701#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg\",\"width\":981,\"height\":501},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=101701#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Scrape The URLs Of A Domain And Write The Results To BigQuery\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=101701","og_locale":"en_US","og_type":"article","og_title":"Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network","og_description":"In my intense love affair with the Google Cloud Platform, I\u2019ve never felt more inspired to write content and try things out. After starting with a Snowplow Analytics setup guide, and continuing with a Lighthouse audit automation tutorial, I\u2019m going to show you yet another cool thing you can do with GCP. In this guide, [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=101701","og_site_name":"Som2ny Network","article_published_time":"2025-02-21T15:20:34+00:00","og_image":[{"width":981,"height":501,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg","type":"image\/jpeg"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"9 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=101701#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=101701"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Scrape The URLs Of A Domain And Write The Results To BigQuery","datePublished":"2025-02-21T15:20:34+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=101701"},"wordCount":1666,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=101701#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg","keywords":["BigQuery","domain","Results","Scrape","URLs","Write"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=101701#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=101701","url":"https:\/\/fivemor.com\/?p=101701","name":"Scrape The URLs Of A Domain And Write The Results To BigQuery - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=101701#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=101701#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg","datePublished":"2025-02-21T15:20:34+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=101701#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=101701"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=101701#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/bigquery-status-codes.jpg","width":981,"height":501},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=101701#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Scrape The URLs Of A Domain And Write The Results To BigQuery"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/101701","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=101701"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/101701\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/101702"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=101701"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=101701"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=101701"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=101701"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=101701"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}