<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom"><generator uri="https://jekyllrb.com/" version="3.8.5">Jekyll</generator><link href="https://refinepro.com/feed.xml" rel="self" type="application/atom+xml"/><link href="https://refinepro.com/" rel="alternate" type="text/html"/><updated>2026-07-30T07:50:05-04:00</updated><id>https://refinepro.com/feed.xml</id><title type="html">RefinePro</title><subtitle>Turn data into competitive advantages</subtitle><entry><title type="html">Download The Data Innovation Canvas</title><link href="https://refinepro.com/blog/download-data-innovation-canvas/" rel="alternate" type="text/html" title="Download The Data Innovation Canvas"/><published>2020-07-09T00:00:00-04:00</published><updated>2020-07-09T00:00:00-04:00</updated><id>https://refinepro.com/blog/download-data-innovation-canvas</id><content type="html" xml:base="https://refinepro.com/blog/download-data-innovation-canvas/">&lt;p&gt;Learn more about the Data Innovation Canvas from Communitech. &lt;/p&gt; &lt;p&gt; Watch Chris Willsher, Director of Data Platforms at Communitech, presents the canvas during the May 8th Communitech® Data Hub Sessions. The video starts at 41:25&lt;/p&gt; &lt;div style=&quot;text-align: center&quot;&gt; &lt;iframe width=&quot;560&quot; height=&quot;315&quot; src=&quot;https://www.youtube.com/embed/3welnQnWDSw?start=2485&quot; frameborder=&quot;0&quot; allow=&quot;accelerometer; autoplay; encrypted-media; gyroscope; picture-in-picture&quot; allowfullscreen=&quot;&quot;&gt;&lt;/iframe&gt; &lt;br /&gt;&lt;br /&gt; &lt;a href=&quot;/images/blog/Data-Innovation-Canvas.pdf&quot; class=&quot;button special mb-2&quot;&gt;Download the data innovation canvas&lt;/a&gt;&lt;/div&gt; &lt;!-- overflow:hidden contains the floated .alignleft image, which is width:100% and taller than the old min-height:250px -- without it the image overflowed the box and the CTA sat under it. --&gt; &lt;div style=&quot;text-align: center; clear:both; overflow:hidden; margin-bottom:3em;&quot;&gt; &lt;a href=&quot;/images/blog/Data-Innovation-Canvas.pdf&quot;&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/data_innovation_canevas.png&quot; width=&quot;100%&quot; /&gt;&lt;/a&gt; &lt;/div&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt;</content><author><name>martin</name></author><summary type="html">Learn more about the Data Innovation Canvas from Communitech. Watch Chris Willsher, Director of Data Platforms at Communitech, presents the canvas during the May 8th Communitech® Data Hub Sessions. The video starts at 41:25</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/data_innovation_canevas.png"/></entry><entry><title type="html">The secret for long term-growth - or why data is the new oil</title><link href="https://refinepro.com/blog/why-data-is-new-oil/" rel="alternate" type="text/html" title="The secret for long term-growth - or why data is the new oil"/><published>2020-07-09T00:00:00-04:00</published><updated>2020-07-09T00:00:00-04:00</updated><id>https://refinepro.com/blog/why-data-is-new-oil</id><content type="html" xml:base="https://refinepro.com/blog/why-data-is-new-oil/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/744407434.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;div&gt;&lt;br /&gt;&lt;/div&gt; &lt;p&gt;Data is the new oil. It stands at the center of an organization’s value proposition, at the core of their product or service creation process. Understanding and managing data is a core competency and not a by-product.&lt;/p&gt; &lt;blockquote&gt; &lt;p&gt;COMPANIES THAT ARE LEADERS IN THE USE OF DATA ARE THREE TIMES MORE LIKELY TO BE FINANCIALLY SUCCESSFUL. source: &lt;a href=&quot;http://www.eiu.com/default.aspx&quot;&gt;Economic Intelligence Unit&lt;/a&gt;&lt;/p&gt; &lt;/blockquote&gt; &lt;p&gt;It’s &lt;a href=&quot;https://www.emc.com/leadership/digital-universe/2012iview/big-data-2020.htm&quot;&gt;estimated&lt;/a&gt; that 40,000 more exabytes of data is either created, replicated, or consumed annually in 2020 compared to 1,200 in 2010. And this increase is happening in almost every industry. Organizations must then learn to exploit and refine data if they want to grow. And to do so, they need to understand how their data strategy evolves at every step of their customer journey and product life cycle. They need their employees to understand how to use, read, understand, and interpret data. And they need to know how to build products and services that are driven by data.&lt;/p&gt; &lt;p&gt;With data, your organization can create products or services that directly help customers. Your offer can take the form of tools (API, data feeds, recommendation engine, etc.) or knowledge and insights (thanks to advanced analytics). But ultimately, data still needs to be at the center of your organization’s business strategy. And the same way the Business Model Canvas is supposed to help organizations develop their business model, Communitech created the &lt;a href=&quot;https://startupheretoronto.com/partners/communitech/communitech-communitech-data-hub-sessions-introducing-the-data-innovation-canvas-2/&quot;&gt;Data Innovation Canvas&lt;/a&gt; to help organizations develop a data model.&lt;/p&gt; &lt;div style=&quot;text-align: center; clear:both;&quot;&gt; &lt;a href=&quot;/blog/download-data-innovation-canvas/&quot; class=&quot;button special mb-2&quot;&gt;Download the Data Innovation Canvas&lt;/a&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/data_innovation_canevas.png&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;p&gt;So you can now start fuelling your organization with data for better growth.&lt;/p&gt; &lt;h3 id=&quot;invest-in-your-future&quot;&gt;INVEST IN YOUR FUTURE&lt;/h3&gt; &lt;p&gt;We all start from an empty (or nearly empty) data store. Building a data strategy is the key to slowly assemble all the data you need, in a structured, efficient, ethical, legal, and reliable way. As you fill up your data warehouse, you can slowly define how your product or your service will incorporate data, and how you can leverage data over time to continually improve your offer.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;It’s a compound interest.&lt;/strong&gt; The longer you invest in your data, i.e. collect, aggregate, enrich, and analyze them, the more value you gain. Over time, you build historical information on specific elements of your service, product, and customer base, that will improve your analysis and allow you to gain new insights. Later down the road, your historical information will help you create better data initiatives. Essentially, you’re being paid off by your own work. You’re making interests over your interests.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Start now&lt;/strong&gt;. Because this history is built over your doing, your innovations, and your products and services, it’s entirely exclusive to you and impossible to reproduce. Yes, it’ll take time. You might not be able to use data to address current challenges. The goal is more in the long term: How can you leverage data that improve &lt;em&gt;over time&lt;/em&gt;? The only way to do so is by thinking two or three steps ahead to set the right goals.&lt;/p&gt; &lt;p&gt;Keep in mind that the more data you collect, the more accurate you’ll be. But the more accurate you’ll be, the more changes you’ll need to make to your products, services, and workflows to incorporate data in your operations.&lt;/p&gt; &lt;p&gt;Data is the new oil, but it’s not sufficient to fill the tank of your machine once and leave for a year. During your data journey, you’ll rethink how to use that fuel efficiently. You will also see new opportunities and new destinations to reach with a bus filled with happy and empowered customers.&lt;/p&gt; &lt;h3 id=&quot;find-the-perfect-productservice-market-fit&quot;&gt;FIND THE PERFECT PRODUCT/SERVICE MARKET FIT&lt;/h3&gt; &lt;p&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/351983435.jpg&quot; width=&quot;45%&quot; /&gt;Once you’ve filled your Data Innovation Canvas, you’ll be ready to build a data-driven business model. In other words, your business model will define what role data will play in the growth of your organization and how it will adapt to data changes.&lt;/p&gt; &lt;p&gt;Adaptability is the keyword here. Our world is moving fast, and that’s already an understatement. An organization that accepts to harness the power of data must also be ready to adapt to its many changes and variations. You should also take into consideration your organization’s learning curve when working with data. Mainly:&lt;/p&gt; &lt;ul&gt; &lt;li&gt; &lt;p&gt;Your business plan should &lt;strong&gt;consider all possible ways of presenting data&lt;/strong&gt; and their potential impact on your revenue model and targeted segments. One dataset, for example, could be used simultaneously to feed different initiatives.&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Data changes all the time&lt;/strong&gt;. It’s constantly being updated, transformed, erased, combined, added, etc. You need to consider the possible change in your data sources overtime, and how it’ll affect your organization and its capacity to deliver its products and services. If you use external data, you need to make sure your architecture is strong and flexible enough to handle unpredictable changes.&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Data is not confined to one department&lt;/strong&gt;. It can, and should, be used by different teams for different applications. The more you use your data, the most reliable it becomes, and the more you find ways to use it!&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Data shouldn’t be complex&lt;/strong&gt;. You don’t need to create the next artificial intelligence unicorn. Start with a narrow scope and explore the real use case for it. If you make it easy to explain, you’ll help increase the transparency of your activities for your users and stakeholders, and thus gain their trust. And this will show to be a lot more efficient than diving headfirst into complex cross-analysis, especially if it’s not needed.&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Your customers change all the time too&lt;/strong&gt;. And their expectations of your organization change also. Keep an eye out where the industry is going.&lt;/p&gt; &lt;/li&gt; &lt;/ul&gt; &lt;p&gt;With a business model based on data, and most importantly based on the volatile nature of data, you multiply your chances of finding the perfect product/service market fit. Instead of relying on time-sensitive market analysis, your growth will be rooted in a deep understanding of your data and different user experience iteration. You’ll drive that bus full of customers not only with the right fuel but, most importantly, in the right direction.&lt;/p&gt; &lt;div style=&quot;text-align: center; clear:both;&quot;&gt; &lt;a href=&quot;/blog/download-data-innovation-canvas/&quot; class=&quot;button special mb-2&quot;&gt;Download the Data Innovation Canvas&lt;/a&gt; &lt;/div&gt; &lt;h3 id=&quot;take-away&quot;&gt;TAKE AWAY&lt;/h3&gt; &lt;p&gt;Now you know: collecting data and learning how to use it is your next big move. But it’s never that easy, isn’t it? There is such a thing as bad data. Before you throw yourself in juicy datasets, make sure you ask yourself the &lt;a href=&quot;/blog/10-questions-to-ask-before-using-new-data/&quot;&gt;10 essential questions&lt;/a&gt; to make sure you’re feeding your organization with relevant data from reliable sources. It’ll also help you make sure the data you’re collecting is usable for your organization, or that you have the required abilities to clean it. And once you’ve clearly defined your source, don’t forget to plan ahead: how will you use these different sources? Can you cross analyze to derive new insights?&lt;/p&gt; &lt;p&gt;Collecting, cleaning, and analyzing data &lt;a href=&quot;/blog/schedule-maintain-web-scraper/&quot;&gt;are not easy&lt;/a&gt; (or cheap) tasks. You need to set the right expectations, or you risk drowning is tasks you hadn’t planned for, driving a bus that’s missing critical pieces of mechanics. It might be easier to start off with data that’s naturally closer to your needs as to limit the cleansing and preparation needed. And then, slowly, as your organization grows, you can improve your data granularity (the level of detail) or linkage (how it relates to other information).&lt;/p&gt; &lt;p&gt;The secret is investing in your infrastructure. Build data foundations from the start, from engineering, based on a solid Data Canvas. Build foundations that will last while you iterate on the front end and insight delivery. Give yourself enough space to package data differently, according to your users.&lt;/p&gt; &lt;p&gt;You’ll find all sorts of “ready to buy” datasets out there. But the datasets that matter are the ones you build. So, &lt;strong&gt;start collecting data early to build historical information. Because one thing you’ll never be able to buy back is time.&lt;/strong&gt;&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="Strategy"/><summary type="html"></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/744407434.jpg"/></entry><entry><title type="html">Download PDF Form 10 Questions Before Using New Data</title><link href="https://refinepro.com/blog/download-10-questions-for-new-data/" rel="alternate" type="text/html" title="Download PDF Form 10 Questions Before Using New Data"/><published>2020-06-15T00:00:00-04:00</published><updated>2020-06-15T00:00:00-04:00</updated><id>https://refinepro.com/blog/Download-Questions-For-New-Data</id><content type="html" xml:base="https://refinepro.com/blog/download-10-questions-for-new-data/">&lt;div style=&quot;text-align: center;&quot;&gt; &lt;img class=&quot;alignnone wp-image-1426 size-medium&quot; src=&quot;/images/blog/10QuestionsBeforeUsingData.png&quot; alt=&quot;PDF preview&quot; width=&quot;225&quot; height=&quot;300&quot; srcset=&quot;/images/blog/10QuestionsBeforeUsingData.png 225w, /images/blog/10QuestionsBeforeUsingData.png 720w&quot; sizes=&quot;(max-width: 225px) 100vw, 225px&quot; /&gt; &lt;/div&gt; &lt;p&gt;An editable PDF form for working through the ten questions on a new dataset. It used to sit behind an email signup; it is now a direct download.&lt;/p&gt; &lt;div style=&quot;text-align:center;&quot;&gt;&lt;a href=&quot;/images/blog/Ten-Questions-to-Ask-Before-Using-New-Data.pdf&quot; class=&quot;button special&quot;&gt;Download the question form (PDF)&lt;/a&gt;&lt;/div&gt; &lt;p&gt;Each question, and why it matters, is covered in &lt;a href=&quot;/blog/10-questions-to-ask-before-using-new-data/&quot;&gt;10 Questions to Ask Before Using New Data&lt;/a&gt;.&lt;/p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt;</content><author><name>martin</name></author><summary type="html"></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/10QuestionsBeforeUsingData.png"/></entry><entry><title type="html">Download PDF Form Web Scraping Software Comparison Table</title><link href="https://refinepro.com/blog/download-web-scraping-comparison-form/" rel="alternate" type="text/html" title="Download PDF Form Web Scraping Software Comparison Table"/><published>2020-06-15T00:00:00-04:00</published><updated>2020-06-15T00:00:00-04:00</updated><id>https://refinepro.com/blog/Download-Web-Scraper-Comparison</id><content type="html" xml:base="https://refinepro.com/blog/download-web-scraping-comparison-form/">&lt;div style=&quot;text-align: center;&quot;&gt; &lt;img class=&quot;alignnone wp-image-1426 size-medium&quot; src=&quot;/images/blog/WebScrapingComparison.png&quot; alt=&quot;PDF preview&quot; height=&quot;300&quot; srcset=&quot;/images/blog/WebScrapingComparison.png 225w, /images/blog/WebScrapingComparison.png 720w&quot; sizes=&quot;(max-width: 225px) 100vw, 225px&quot; /&gt; &lt;/div&gt; &lt;p&gt;An editable PDF form for comparing three web scraping solutions side by side. It used to sit behind an email signup; it is now a direct download.&lt;/p&gt; &lt;div style=&quot;text-align:center;&quot;&gt;&lt;a href=&quot;/images/blog/Web-Scrapring-Software-Comparison-Table.pdf&quot; class=&quot;button special&quot;&gt;Download the comparison table (PDF)&lt;/a&gt;&lt;/div&gt; &lt;p&gt;The criteria, and the reasoning behind choosing on criteria rather than benchmarks, are in &lt;a href=&quot;/blog/who-why-why-of-web-scraping/&quot;&gt;Who, Why and What of Web Scraping&lt;/a&gt;.&lt;/p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt;</content><author><name>martin</name></author><summary type="html"></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/WebScrapingComparison.png"/></entry><entry><title type="html">10 questions to ask before using new data</title><link href="https://refinepro.com/blog/10-questions-to-ask-before-using-new-data/" rel="alternate" type="text/html" title="10 questions to ask before using new data"/><published>2020-05-25T00:00:00-04:00</published><updated>2020-05-25T00:00:00-04:00</updated><id>https://refinepro.com/blog/10-questions-to-ask-before-using-new-data.md</id><content type="html" xml:base="https://refinepro.com/blog/10-questions-to-ask-before-using-new-data/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1496369480.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;p&gt;Data extraction projects are complex and often require quite a lot of time and effort. To make sure your organization is creating value and that your money and your time are well spent, the first logical step is to choose your sources carefully. To help you achieve just that, we create &lt;strong&gt;a list of 10 questions you need to ask before you set your sights on a dataset&lt;/strong&gt;. The goal here is to collect and analyze all the data existing information in order to clarify its ownership, publication, structure, content, quality, relationship, etc. Only by going through this process can you guarantee the suitability of your sources and identify potential problems and particularities.&lt;/p&gt; &lt;p&gt;This checklist will help you assess all the elements you need to know in order to proceed with your data project. Most of all, once you have all the answers, you will have everything you need to define what will be your game plan to transform and manipulate the datasets you chose.&lt;/p&gt; &lt;p&gt;So, without further ado, here are &lt;strong&gt;ten questions to ask before using new data.&lt;/strong&gt;&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Question 1: Who owns the data?&lt;/li&gt; &lt;li&gt;Question 2: Who publishes the data?&lt;/li&gt; &lt;li&gt;Question 3: Is the dataset documented?&lt;/li&gt; &lt;li&gt;Question 4: How is the data collected?&lt;/li&gt; &lt;li&gt;Question 5: How is the data maintained and updated?&lt;/li&gt; &lt;li&gt;Question 6: What are the format and granularity?&lt;/li&gt; &lt;li&gt;Question 7: Does the data follow standards?&lt;/li&gt; &lt;li&gt;Question 8: Can you link your data to another dataset?&lt;/li&gt; &lt;li&gt;Question 9: Under what licenses the data is released?&lt;/li&gt; &lt;li&gt;Question 10: Are there data privacy issues?&lt;/li&gt; &lt;/ul&gt; &lt;div style=&quot;text-align:center;&quot;&gt; &lt;a href=&quot;/blog/download-10-questions-for-new-data/&quot; class=&quot;button special&quot;&gt;Download our editable form for your personal use&lt;/a&gt;&lt;/div&gt; &lt;p&gt;&lt;br /&gt;&lt;/p&gt; &lt;h3 id=&quot;question-1-who-owns-the-data&quot;&gt;Question 1. Who owns the data?&lt;/h3&gt; &lt;p&gt;And by “own,” we don’t necessarily mean “publish” (see question 2). You need to know where the data originally comes from, and whom to contact if you ever have any questions or issues that need solving. Also, if you ever need to attribute ownership (see question 9) when reusing the data, this owner will be the one you will refer too. Basically, you need to put a human face and a name to the data you wish to extract and use.&lt;/p&gt; &lt;h3 id=&quot;question-2-who-publishes-the-data&quot;&gt;Question 2. Who publishes the data?&lt;/h3&gt; &lt;p&gt;There are a lot of platforms out there that offer huge datasets, like &lt;a href=&quot;https://www.quandl.com/&quot;&gt;Quandl&lt;/a&gt;, and &lt;a href=&quot;https://data.world/&quot;&gt;data.world&lt;/a&gt;. But it doesn’t mean they own the data they share with you. It is, therefore, imperative that you know who owns and publishes and shares the data. This distinction will help you better answer the following questions.&lt;/p&gt; &lt;h3 id=&quot;question-3-is-the-dataset-documented&quot;&gt;Question 3. Is the dataset documented?&lt;/h3&gt; &lt;p&gt;You need to gather all the information on how the data was collected (see questions 4 and 5) and how it should be interpreted (see question 7). This includes the schema of the data, with the data type and validation rules. The goal here is to make sure you can answer most questions without having to go back to the data owner.&lt;/p&gt; &lt;h3 id=&quot;question-4-how-is-the-data-collected&quot;&gt;Question 4. How is the data collected?&lt;/h3&gt; &lt;p&gt;The answer to this question will help you identify any potential biases in the way data is collected. It will also give you some extremely important information concerning the data itself: is the data complete or partial? Has it been pre-processed before its publication? You want to know what the original state of the data was and how much it has changed (or not) before reaching you.&lt;/p&gt; &lt;h3 id=&quot;question-5-how-is-the-data-maintained-and-updated&quot;&gt;Question 5. How is the data maintained and updated?&lt;/h3&gt; &lt;p&gt;Now that you know what the data looked like originally, you want to know what processes it goes through. For example, you want to ensure your data will still be reliable in the long-term. You also want to know how often it’s updated and if the set contains all the records or only updated ones. You also need to know if there’s a change in the collect methodology, or if the dataset stops being available.&lt;/p&gt; &lt;p&gt;Without these answers, you might end up building a script for something that won’t be available in two days’ time, or not in the format you expected.&lt;/p&gt; &lt;h3 id=&quot;question-6-what-are-the-format-and-granularity&quot;&gt;Question 6. What are the format and granularity?&lt;/h3&gt; &lt;p&gt;You must identify the formats in which your data is made available. Formats can usually be categorized as follows:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;NonFriendly Format (PDF, web page, Word document, image)&lt;/li&gt; &lt;li&gt;Flat file (CSV, XLS)&lt;/li&gt; &lt;li&gt;Structured file (JSON, XML)&lt;/li&gt; &lt;li&gt;API and web service, provided by the source or by a third party (Quandl, data.world)&lt;/li&gt; &lt;li&gt;Maps (KML, Shapefile, GeoJSON)&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;Once you know what format you’ll need to deal with, you can better choose the tools and solutions you’ll need to extract and transform your data. It will also help you identify the &lt;strong&gt;data granularity, or the lowest data point available&lt;/strong&gt;. If, for example, your information concerns time, you want to know if the smallest possible data point is second, minute, hour, day, month, or year. The same goes for maps: address, postal code, city, or state?&lt;/p&gt; &lt;p&gt;Knowing what your data looks like and what shape it takes will guarantee that you have all the information you need to extract it using the right tools and solutions.&lt;/p&gt; &lt;h3 id=&quot;question-7-were-specific-standards-applied-to-the-dataset&quot;&gt;Question 7. Were specific standards applied to the dataset?&lt;/h3&gt; &lt;p&gt;When data is collected and published according to certain standards, it helps remove ambiguity on the collection, aggregation, and preparation methods. Standardization also allows us to compare and combine data according to jurisdiction or period. Data regarding elections, &lt;a href=&quot;http://open311.org/&quot;&gt;311 calls&lt;/a&gt;, census or &lt;a href=&quot;https://developers.google.com/transit/gtfs&quot;&gt;transit&lt;/a&gt; information, for example, are all standardized.&lt;/p&gt; &lt;h3 id=&quot;question-8-can-you-link-the-data-to-another-dataset&quot;&gt;Question 8. Can you link the data to another dataset?&lt;/h3&gt; &lt;p&gt;When profiling your data, you need to make sure you understand all its relationships with other datasets. Is it isolated, or could it be combined with other internal or external data? What new insight can you build from it? How could you merge them? Do they share a common key? Basically, the goal here is to profile your data as it related to other data.&lt;/p&gt; &lt;h3 id=&quot;question-9-was-the-data-published-under-a-license-and-if-so-which-one&quot;&gt;Question 9. Was the data published under a license? And if so, which one?&lt;/h3&gt; &lt;p&gt;Some organizations will choose to publish their data under a license, which then defines how the data should be collected, shared, and used. You need to be aware of these licenses and understand how they work. The most common are:&lt;/p&gt; &lt;h4 id=&quot;odc-public-domain-dedication-and-licence-pddl&quot;&gt;ODC Public Domain Dedication and Licence (PDDL)&lt;/h4&gt; &lt;div style=&quot;clear:both;&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/share.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/remix.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/pd.svg&quot; width=&quot;7%&quot; /&gt;&lt;/div&gt; &lt;p&gt;With &lt;a href=&quot;https://www.opendatacommons.org/licenses/pddl/1-0/index.html&quot;&gt;PPDL&lt;/a&gt;, users can share, create, and adapt the document. There’s no restriction and the dataset is public domain.&lt;/p&gt; &lt;h4 id=&quot;open-data-commons-attribution-license-odc-by&quot;&gt;Open Data Commons Attribution License (ODC-By)&lt;/h4&gt; &lt;div style=&quot;clear:both;&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/share.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/remix.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/by.svg&quot; width=&quot;7%&quot; /&gt;&lt;/div&gt; &lt;p&gt;With &lt;a href=&quot;https://opendatacommons.org/licenses/by/index.html&quot;&gt;ODC-By&lt;/a&gt;, users can share, create, and adapt the document. The only restriction is the attribution, which means that users need to cite the source.&lt;/p&gt; &lt;h4 id=&quot;open-data-commons-open-database-license-odc-odbl&quot;&gt;Open Data Commons Open Database License (ODC-ODbL)&lt;/h4&gt; &lt;div style=&quot;clear:both;&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/share.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/remix.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/by.svg&quot; width=&quot;7%&quot; /&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/sa.svg&quot; width=&quot;7%&quot; /&gt;&lt;/div&gt; &lt;p&gt;With &lt;a href=&quot;https://opendatacommons.org/licenses/odbl/index.html&quot;&gt;ODC-ODbL&lt;/a&gt;, users can share, create, and adapt the document, but they need to cite the sources, share with the same license.&lt;/p&gt; &lt;h4 id=&quot;custom-licenses&quot;&gt;Custom licenses&lt;/h4&gt; &lt;div style=&quot;clear:both;&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/custom-license.png&quot; width=&quot;7%&quot; /&gt;&lt;/div&gt; &lt;p&gt;Unfortunately, custom licenses are extremely popular. They necessitate that you read it, understand it, and make sure you respect its specific requirements. This could impact how you can collect and transform your data, as well as how you can use it.&lt;/p&gt; &lt;h3 id=&quot;question-10-are-there-data-privacy-issues&quot;&gt;Question 10. Are there data privacy issues?&lt;/h3&gt; &lt;p&gt;Privacy is protected differently from state to state. Important differences even exist between Canada, the United States and Europe. You need to know if your datasets contain Personally Identifiable Information (PII) or if it would be possible to &lt;a href=&quot;https://georgetownlawtechreview.org/re-identification-of-anonymized-data/GLTR-04-2017/&quot;&gt;re-identify individuals based on anonymized data&lt;/a&gt;. This is especially true with &lt;a href=&quot;https://www.ncbi.nlm.nih.gov/books/NBK208613/&quot;&gt;healthcare data&lt;/a&gt;.&lt;/p&gt; &lt;div style=&quot;text-align:center;&quot;&gt; &lt;a href=&quot;/blog/download-10-questions-for-new-data/&quot; class=&quot;button special&quot;&gt;Download our form to make your own analysis&lt;/a&gt;&lt;/div&gt; &lt;p&gt;&lt;br /&gt;&lt;/p&gt; &lt;h3 id=&quot;and-with-that&quot;&gt;AND WITH THAT&lt;/h3&gt; &lt;p&gt;Knowing what data you need is not enough to start a data extraction project. More than “what,” you need to know “who” your data is. Knowing your data is the only way to know for sure you’re using the right tool, at the right schedule, with the right script, to get the right data and transform it correctly.&lt;/p&gt; &lt;p&gt;This list of ten questions should be your first step in defining if a dataset is worth all the efforts you’re ready to put into it. Data projects are complex projects on their own and they require that you plan them well.&lt;/p&gt; &lt;p&gt;Choose your sources is only the first step to a long story. Depending on your sources and needs (are you dealing with unfriendly &lt;a href=&quot;/expertise/pdf-extraction/&quot;&gt;formats like PDFs?&lt;/a&gt;, you’ll need to define the best tools for &lt;a href=&quot;/expertise/web-scraping/&quot;&gt;web scraping&lt;/a&gt;, the best way to &lt;a href=&quot;/blog/how-to-maintain-data-quality/&quot;&gt;maintain data quality&lt;/a&gt; throughout the whole process, the best way to &lt;a href=&quot;/blog/14-rules-for-successful-ETL/&quot;&gt;build a solid ETL process&lt;/a&gt;, and the &lt;a href=&quot;/expertise/design-architecture/&quot;&gt;best architecture for data extraction processes.&lt;/a&gt;&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="Strategy"/><category term="Web Scraping"/><summary type="html">Data extraction projects are complex and often require quite a lot of time and effort. To make sure your organization is creating value and that your money and your time are well spent, the first logical step is to choose your sources carefully. To help you achieve just that, we create a list of 10 questions you need to ask before you set your sights on a dataset. The goal here is to collect and analyze all the data existing information in order to clarify its ownership, publication, structure, content, quality, relationship, etc. Only by going through this process can you guarantee the suitability of your sources and identify potential problems and particularities.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/1496369480.jpg"/></entry><entry><title type="html">PDF extraction - Everything you need to know</title><link href="https://refinepro.com/blog/what-to-know-about-pdf-extraction/" rel="alternate" type="text/html" title="PDF extraction - Everything you need to know"/><published>2020-05-18T00:00:00-04:00</published><updated>2020-05-18T00:00:00-04:00</updated><id>https://refinepro.com/blog/PDF-extraction-Everything-you-need-to-know</id><content type="html" xml:base="https://refinepro.com/blog/what-to-know-about-pdf-extraction/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1505621234.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;p&gt;Our team here at RefinePro has a deep experience doing research and development in data processing and automation. And PDF extraction is one of the many services we offer.&lt;/p&gt; &lt;p&gt;But before we go into too much detail … &lt;strong&gt;why exactly do people need PDF extraction for?&lt;/strong&gt;&lt;/p&gt; &lt;p&gt;Portable Document Format, or PDF, is a standardized file format. It allows users to distribute read-only documents that will present the same text and images independently of the hardware, software, or operating system used to open it (Mac, Windows, Linux, iPhone, Android, and others). PDF documents may contain a wide variety of information other than text and graphics, such as interactive elements (annotations and editable fields), structural elements, media, and various other content formats.&lt;/p&gt; &lt;p&gt;In today’s work environment, PDF is often the go-to solution for exchanging business data. Suppliers, for example, mostly prefer PDF to create their price lists and catalogues and to exchange invoices, purchase orders, reports, etc. So, whether you’re trying to gather a larger volume of &lt;strong&gt;data on a specific subject&lt;/strong&gt; in your field of research or just trying to extract a list of items and prices for your &lt;strong&gt;eCommerce website&lt;/strong&gt;, you need to find a way to convert information contained in PDF documents into usable structured data.&lt;/p&gt; &lt;p&gt;And let’s be honest, nobody wants to (or can!) go through doze or even hundreds of documents manually.&lt;/p&gt; &lt;p&gt;PDF documents are easy to read for humans, but they rarely contain any machine-readable data. Their format varies considerably from one file to another, depending on how it was generated. If you’re lucky, the document you’re extracting your data from is in text format, with numbers organized neatly in tables. But if you’re not lucky, the information is embedded in an image. In that case, you’ll need to use Optical Character Recognition (OCR) to help you get the data.&lt;/p&gt; &lt;p&gt;Accessing a massive amount of information stored in PDFs and converting it can then be a burdensome task. Luckily, PDF data extraction offers solutions to automate this task and automatically convert messy information into structured and usable data. And PDF extraction projects are no news for us. We invested in some of the proven technologies, and we are always testing out new software to make sure we help you build the data extraction project you need to meet your goals.&lt;/p&gt; &lt;h3 id=&quot;1-pdf-extraction-how&quot;&gt;1. PDF EXTRACTION: HOW?&lt;/h3&gt; &lt;h4 id=&quot;11-the-right-tool-for-your-project&quot;&gt;1.1 THE RIGHT TOOL FOR YOUR PROJECT&lt;/h4&gt; &lt;p&gt;There are a lot of different systems out there to help you set a solid PDF extraction project. For business analysts, it’s often easier to go with “What You See Is What You Get” interfaces (&lt;strong&gt;WYSIWYG&lt;/strong&gt;) like &lt;a href=&quot;https://docparser.com/?ref=roqts&quot;&gt;DocParser&lt;/a&gt;. These systems tend to be more expensive, but they are easy to use and set, and they work well with high volume of easy cases. For entry-level programmers, some solutions offer more flexibility and low code complexity, which makes it easier to support exceptions for complex files. However, they still require programming knowledge and expertise on data extraction project as a whole. They usually run on &lt;strong&gt;JAVA&lt;/strong&gt; or &lt;strong&gt;Python&lt;/strong&gt;.&lt;/p&gt; &lt;p&gt;The advantage of working with a partner like RefinePro is that thanks to our years of experience, we can help you &lt;strong&gt;select the technology that will best answer your requirements.&lt;/strong&gt; We listed the four categories you should keep in mind.&lt;/p&gt; &lt;h4 id=&quot;12-assessing-your-needs&quot;&gt;1.2 ASSESSING YOUR NEEDS&lt;/h4&gt; &lt;p&gt;&lt;strong&gt;Your business and legal requirements:&lt;/strong&gt; You should ask yourself:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Are you working with sensitive data? What privacy laws do you need to comply with?&lt;/li&gt; &lt;li&gt;Do you want to use non-open source technology?&lt;/li&gt; &lt;li&gt;What level of dependency do you want or can have on a service or technology provider?&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;&lt;strong&gt;The connectivity to your systems:&lt;/strong&gt; This includes the method used to send and receive the PDFs with your systems (e.g. via an API, a database connection, or other) and if you want to process files in batch or on-demand as they are collected?&lt;/p&gt; &lt;p&gt;&lt;strong&gt;The volume of data:&lt;/strong&gt; including how many are your processing per day? How many different layouts? What are the data validation rules (schema, business rules, etc.); and what happens when the validation job rejects data (the review process).&lt;/p&gt; &lt;p&gt;&lt;strong&gt;Your Resources:&lt;/strong&gt; Who will monitor your PDF extraction project? What type of skills (and training) do they need? What kind of medium- and long-term support do you need?&lt;/p&gt; &lt;h3 id=&quot;2-refinepros-pdf-extraction-subsystems&quot;&gt;2. REFINEPRO’S PDF EXTRACTION SUBSYSTEMS&lt;/h3&gt; &lt;p&gt;Over the years, we have developed an extraction architecture that relies on a set of best practices and proven engineered patterns. We recommend &lt;a href=&quot;/blog/divide-and-conquer-your-data-project/&quot;&gt;decoupling your steps&lt;/a&gt; to make troubleshooting easier. PDF extraction should follow four steps: data collection, data normalization, data validation, and delivery.&lt;/p&gt; &lt;p&gt;These steps are part of an architecture in which ingestion and normalization of each PDF document are divided into three subsystems.&lt;/p&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/pdfprocess.png&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;h4 id=&quot;21-subsystem-one-collection-and-normalization&quot;&gt;2.1 Subsystem One. Collection and Normalization.&lt;/h4&gt; &lt;p&gt;In the first part, we bring together collection and normalization. All the different formats of data collected are being morphed into a standard schema, which is the set of validation rules you implemented to define what a “good” data is. To do so, the developer writes one PDF extraction and one normalization script per PDF layout. In other words, different scripts are used depending on the outlines, style, and logical component content of the PDF. This way, one script will extract data from documents matching the same layout—the same logical structure—to then transform it in a usable format for your team.&lt;/p&gt; &lt;p&gt;In this script, the developer will add all the exceptions related to a specific PDF layout so that each file format can be processed independently. This way, if one script returns an error, it only affects one layout and not the entire project, making troubleshooting easier.&lt;/p&gt; &lt;p&gt;For the normalization step, more specifically, &lt;a href=&quot;/toolbox/openrefine/&quot;&gt;OpenRefine&lt;/a&gt; is a great tool if you want to build a fully WYSIWYG solution (something we can help you with). On the other hand, &lt;a href=&quot;/toolbox/talend/&quot;&gt;Talend Open Studio&lt;/a&gt;, is perfect if you want to outsource the work to entry-level programmers. We can also &lt;a href=&quot;/offering/training/&quot;&gt;train your team&lt;/a&gt; and help you launch your first project!&lt;/p&gt; &lt;h4 id=&quot;22-subsystem-two-validation-and-delivery-or-the-delivery-of-quality-data&quot;&gt;2.2 Subsystem Two. Validation and delivery (or the delivery of quality data)&lt;/h4&gt; &lt;p&gt;During the second part, which includes validation and delivery, we leverage a unified schema. We only need one validation and one delivery script for all PDF layout. The data is being validated using the schema to ensure compliance with your business rules before it is delivered into your system.&lt;/p&gt; &lt;p&gt;During validation, we define and document the schema, namely the elements that make a “good” data. As such, a validation error occurs when an extracted data doesn’t pass the validation rules established for the project. This corruption can come from a bug in the workflow, or changes in the data sources.&lt;/p&gt; &lt;p&gt;This step is particularly important. When we develop a PDF extraction project script, one of the priorities is to create a validation script to ensure we do not over-engineered data quality. We need to ensure that the validation steps fail as early as possible to avoid corrupting downstream systems.&lt;/p&gt; &lt;h4 id=&quot;23-subsystem-three-scheduling-monitoring-and-maintaining&quot;&gt;2.3 Subsystem Three. Scheduling, Monitoring, and Maintaining&lt;/h4&gt; &lt;p&gt;The third part is the use of infrastructure or platform to execute, schedule, configure, and monitor the scripts themselves to ensure they keep delivering reliable data. Most importantly, also, data quality (article 3) will need to be monitored thoroughly.&lt;/p&gt; &lt;h3 id=&quot;what-about-ai&quot;&gt;WHAT ABOUT AI?&lt;/h3&gt; &lt;p&gt;Artificial Intelligence is the new kid on the block. Everyone knows it, everybody wants to use it, many people claim to have mastered it, but few people actually offer it. In PDF extraction, more specifically, we have seen a lot of promising development, but we’re not there yet. AI can be used for very narrow use cases. Instead of trying to find the next shiny object, we recommend sticking to well-proven and tested solutions that will help you get the results you’re looking for. Be sure, however, that our team is keeping a close eye on all the new technologies out there. Don’t hesitate to contact us curious@refinepro.com if you’d like an independent assessment on a specific software.&lt;/p&gt; &lt;h3 id=&quot;and-with-that&quot;&gt;AND WITH THAT&lt;/h3&gt; &lt;p&gt;Here at RefinePro, we provide data strategy, system architecture, implementation, and outsourcing services to help organizations scale and automate data acquisition and transformation workflows. Whether you decide to work with us, with another service provider, or even on your own, you’ll need to make sure to select the right tools (and not just the PDF extracting tool: database, servers, data processing framework, etc.) and set up your processes to meet your data quality requirements while minimizing the maintenance efforts.&lt;/p&gt; &lt;p&gt;For years, we have helped clients define what system and process to put in place to ensure their needs are answered in the most time- and cost-efficient manner. So, before you throw yourself on Google or your in-house expertise to develop a complex data extraction project, contact us!&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="PDF"/><summary type="html">Our team here at RefinePro has a deep experience doing research and development in data processing and automation. And PDF extraction is one of the many services we offer.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/1505621234.jpg"/></entry><entry><title type="html">How to divide and conquer your data project for success</title><link href="https://refinepro.com/blog/divide-and-conquer-your-data-project/" rel="alternate" type="text/html" title="How to divide and conquer your data project for success"/><published>2020-05-17T00:00:00-04:00</published><updated>2020-05-17T00:00:00-04:00</updated><id>https://refinepro.com/blog/divide-and-conquer-your-data-project</id><content type="html" xml:base="https://refinepro.com/blog/divide-and-conquer-your-data-project/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1257325042.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;p&gt;Data extraction is now one of the most efficient ways for companies to stay up to date with current events and trends, but also to position themselves in their field. But for a lot of small entrepreneurs and even larger companies, the implementation of data extraction projects presents new challenges: How should these processes be implemented, and by whom?&lt;/p&gt; &lt;p&gt;Web Scraping is known as the process by which data is extracted from different sources and then transformed into usable information. As such, a huge part of any web scraping project relies on a strong Extraction, Transformation, and Loading process, known as ETL. But building a solid ETL architecture for your web scraping requires a lot of technical know-how, combined with the knowledge necessary to adapt these “easy-to-use” tools to your specific needs. Most importantly, your project will also rely on many other crucial processes, including data quality management and administrative procedures.&lt;/p&gt; &lt;p&gt;In this article, we explain &lt;strong&gt;why all your different data extraction processes should be decoupled for a more seamless workflow.&lt;/strong&gt; That might sound counterintuitive… But the idea is as old as the world: divide and conquer (even algorithms understand). Divide your script, divide your task, find solution to sub-problems instead of a major crash, and get reliable end results, without draining your economic and human resources.&lt;/p&gt; &lt;h3 id=&quot;1-what-is-etl&quot;&gt;1. WHAT IS ETL?&lt;/h3&gt; &lt;p&gt;During ETL, data is being copied from pre-defined sources before reaching you in a format that makes it usable. An ETL developer can help you build an architecture that will support the ETL process of your project.&lt;/p&gt; &lt;p&gt;Why an “ETL” developer? A developer creating a robust data transformation process work at the crossroad of different fields and execute functions as diversified as database analysis, system integration, and data transformation development. He or she must ensure that every aspect of the data life cycle has been addressed to ensure its operability and maintainability.&lt;/p&gt; &lt;h3 id=&quot;2-the-extraction-transformation-and-loading-behind-etl&quot;&gt;2. THE EXTRACTION, TRANSFORMATION, AND LOADING BEHIND ETL&lt;/h3&gt; &lt;p&gt;The three main steps of ETL will come as no surprise: extraction, transformation, and loading. But as one might suspect, each of these steps hides a lot of sub-steps that need to be considered. Data going through an ETL process will undergo different stages in its journey. We will now examine how these stages integrate within the three main steps of ETL. And keep in mind our &lt;em&gt;“divide to conquer”&lt;/em&gt; motto: every step of your ETL as its own logic and should be considered separately.&lt;/p&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/dataflow.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;h3 id=&quot;21-discovery-and-curation&quot;&gt;2.1 Discovery and curation&lt;/h3&gt; &lt;p&gt;Before you sit down with your developer and start coding, you must define precisely what sources you want to extract your data from and how you’re going to document all the relevant information (including the owner, the availability, the updates, etc.). At the end,&lt;/p&gt; &lt;ul&gt; &lt;li&gt;You know who published your data, when, and how;&lt;/li&gt; &lt;li&gt;You know what your data looks like (its size, format, relations to other data), and;&lt;/li&gt; &lt;li&gt;You have a map of your data, from the source to your database.&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;Once discovery and curation are done, you can now jump on step two, data collection.&lt;/p&gt; &lt;h3 id=&quot;22-data-collection-extract&quot;&gt;2.2 Data collection (Extract)&lt;/h3&gt; &lt;p&gt;At this substep, your focus should be on getting to data out from its original format. You will need to select the right data extraction technology, whether it is to get data from &lt;a href=&quot;/expertise/migration-integration/&quot;&gt;another database&lt;/a&gt;, XLS files, a &lt;a href=&quot;/expertise/web-scraping/&quot;&gt;website&lt;/a&gt; or a &lt;a href=&quot;/expertise/pdf-extraction/&quot;&gt;PDF document&lt;/a&gt;. Once it has been extracted, data can now be moved to a landing database in your dedicated data transformation environment. It’s from this new environment, entirely under your control, that you can start reviewing and troubleshooting the data.&lt;/p&gt; &lt;h3 id=&quot;23-normalization-and-validation-transform&quot;&gt;2.3 Normalization and validation (Transform)&lt;/h3&gt; &lt;p&gt;Now that you’ve extracted your data, you need to transform it. During the normalization and validation, the messy data you obtained is being prepared to match the format of your target system, whether it’s a data warehouse or your eCommerce website. Before you start developing a transformation process, though, make sure you know what the &lt;a href=&quot;/blog/14-rules-for-successful-ETL&quot;&gt;best practices&lt;/a&gt; are and which ones you need to implement (and ignore).&lt;/p&gt; &lt;p&gt;But remember! Divide to conquer: data extraction (step 1) should be decoupled from the transformation (step 2). Why?&lt;/p&gt; &lt;ul&gt; &lt;li&gt;It makes debugging and restartability easier by segmenting more precisely the journey of your data (see our &lt;a href=&quot;/blog/how-to-maintain-data-quality/&quot;&gt;article on data quality&lt;/a&gt;).&lt;/li&gt; &lt;li&gt;It gives your (ETL) developer the possibility to select the best tool for each job. This way, your web scraper will do its job of scraping (which is already a &lt;a href=&quot;/blog/schedule-maintain-web-scraper/&quot;&gt;complex job in itself&lt;/a&gt;) and your data cleansing tool will do its job of cleaning.&lt;/li&gt; &lt;/ul&gt; &lt;h3 id=&quot;24-enrichment-and-processing-transform&quot;&gt;2.4 Enrichment and processing (Transform)&lt;/h3&gt; &lt;p&gt;Enrichment and processing steps add value to your data by connecting it with other datasets, such as your own proprietary data (like your customer or product information), for example, or other collected datasets. At this stage, you add your business logic, your secret sauce, so that your team can read and make sense of the data. By adding your business logic to messy data, you give yourself the possibility to develop real business intelligence. And that’s when data extraction really becomes interesting.&lt;/p&gt; &lt;p&gt;Again, decoupling is important here. You should have a single processing script for all your data sources. It is a good time to create historical values by comparing records, a process often referred to as the &lt;a href=&quot;https://en.wikipedia.org/wiki/Slowly_changing_dimension&quot;&gt;six slow changing dimensions&lt;/a&gt; (mainly the ability to keep track of unpredictable changes). Building historical information from extracted data gives you an edge (if not an unfair advantage) in understanding your market and industry. You could, for example, track the price of an item over time to predict when it will be on sale or out of stock.&lt;/p&gt; &lt;h3 id=&quot;25-delivery-and-consumption-load&quot;&gt;2.5 Delivery and consumption (Load)&lt;/h3&gt; &lt;p&gt;Your data is now ready to be used. But in order to do so, you need to move it from your external database into a data warehouse available to your team. During this last substep, we read data from the staging or landing table and insert it into your system. We can push it into a warehouse, upload it directly into your operational system (like an eCommerce website), or make it available to your organization via a custom API.&lt;/p&gt; &lt;h3 id=&quot;26-administration-making-it-all-work-together&quot;&gt;2.6 Administration, making it all work together&lt;/h3&gt; &lt;p&gt;It’s the long-forgotten process, but still a crucial one. The administration will orchestrate the many processes we previously covered. A well-administered project will manage depencies between steps, and make sure everything happens in the right order. Your main tools, including your web scraper, will need to be well &lt;a href=&quot;/blog/schedule-maintain-web-scraper/&quot;&gt;scheduled, maintained, and monitored&lt;/a&gt;. It will also be crucial that you implement practices such as &lt;a href=&quot;/blog/14-rules-for-successful-ETL&quot;&gt;logging, code management, configuration, and project management&lt;/a&gt;.&lt;/p&gt; &lt;p&gt;Data discovery and curation, data collection, normalization and validation, enrichment and processing, delivery and consumption, and administration: these are the multiple layers you find when your start scratching the varnish of a data extraction project. They all contribute to the overall stability, but their uniqueness is also what makes the whole structure stronger. With a well-built architecture, the whole system is not affected if one layer is posing a problem or needs to be upgraded for any technical or business reason.&lt;/p&gt; &lt;h3 id=&quot;3-and-with-that&quot;&gt;3. AND WITH THAT&lt;/h3&gt; &lt;p&gt;Obviously, a complex process such as this can’t be built in a day. It takes trial and error to select the best technology and processes, and also to confirm that you can build value from the data. Every data project has its own specificities. Your organization may have a short-term, extremely precise need for data extraction, but most companies want data that support their business strategy, one they can rely on in the long term to make important decisions. This is why we recommend following an &lt;a href=&quot;/blog/agile-data-process/&quot;&gt;agile data transformation process&lt;/a&gt;. We suggest developing complex data products by starting small with a couple of sources before scaling it to a robust business-wide data factory.&lt;/p&gt; &lt;p&gt;Over the years, we’ve built a strong experience in developing and managing these different processes. Our clients rely on us to manage every aspect, layer, and step of their data collection project so they can focus on building lasting insight and product. By doing so, they know they’re paying the right price for their data, and that they’re not overwhelming their development team.&lt;/p&gt; &lt;p&gt;Depending on your team know-how level, your needs will also change. We can deliver custom training for your team before they throw themselves into web scraping, help you build your project from scratch, or even assist you after its implementation. We offer &lt;a href=&quot;/offering/training/&quot;&gt;training and mentoring&lt;/a&gt; as well as &lt;a href=&quot;/offering/team-augmentation/&quot;&gt;team augmentation&lt;/a&gt;, and &lt;a href=&quot;/offering/team-platform/&quot;&gt;data-first application development&lt;/a&gt;.&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="Orchestration"/><category term="ETL"/><category term="Web Scraping"/><summary type="html">Data extraction is now one of the most efficient ways for companies to stay up to date with current events and trends, but also to position themselves in their field. But for a lot of small entrepreneurs and even larger companies, the implementation of data extraction projects presents new challenges: How should these processes be implemented, and by whom?</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/1257325042.jpg"/></entry><entry><title type="html">14 rules to succeed with your ETL project</title><link href="https://refinepro.com/blog/14-rules-for-successful-ETL/" rel="alternate" type="text/html" title="14 rules to succeed with your ETL project"/><published>2020-05-15T00:00:00-04:00</published><updated>2020-05-15T00:00:00-04:00</updated><id>https://refinepro.com/blog/14-rules-to-succeed-with-your-ETL-project</id><content type="html" xml:base="https://refinepro.com/blog/14-rules-for-successful-ETL/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1238032837.jpg&quot; width=&quot;100%&quot; /&gt;&lt;/div&gt; &lt;p&gt;Extracting, transforming, and loading (ETL) data is a complex process at the center of most organizations’ data extraction projects. As we saw in our article on &lt;a href=&quot;/blog/divide-and-conquer-your-data-project&quot;&gt;web scraping and ETL&lt;/a&gt;, the implementation of an ETL workflow is a process that requires a lot of in-depth knowledge in several subfields of statistics and programming.&lt;/p&gt; &lt;p&gt;ETL developers thus work at the crossroad of different fields. They must ensure that every aspect of the data life cycle has been addressed to ensure its operability and maintainability. To help your developer navigate the deep and dark waters of ETL, we’ve drawn on our years of experience to create a list of ETL principles and best practices.&lt;/p&gt; &lt;p&gt;What you see here is not meant to be a grocery list; these guidelines need to be considered, and then implemented or rejected. Your developer will draw on their understanding of the project and their experience to decide which principles are needed, when, and at what range.&lt;/p&gt; &lt;h3 id=&quot;ten-best-practices-for-etl-workflow-implementation&quot;&gt;TEN BEST PRACTICES FOR ETL WORKFLOW IMPLEMENTATION&lt;/h3&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/16.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 1. Modularity&lt;/h4&gt; &lt;p&gt;Modularity is the process of writing reusable code structures to help you keep your job consistent in terms of size and functionalities. With modularity, your project structure is easier to understand, making troubleshooting easier too. The ultimate goal is to improve job readability and maintainability by avoiding the need to write the same code over and over again.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/14.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 2. Atomicity&lt;/h4&gt; &lt;p&gt;Atomicity is used to break down complex jobs into independent and more understandable parts. The workflow is divided between distinct units of work and small and individual executable processes. Each of these parts can be executed separately. It makes testing and troubleshooting easier since the developer doesn’t need to run a long-running process to debug a single operation.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/19.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 3. Change Detection and Increment&lt;/h4&gt; &lt;p&gt;A change detection strategy detects differences and allows incremental data loading. This means that only records changed since the last update is brought into the ETL process, avoiding unnecessary transformations.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/17.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 4. Scalability&lt;/h4&gt; &lt;p&gt;Your ETL process should be implemented in a way that ensures its scalability to allow your project to adapt to the growing volume of data. This way, you don’t have to redefine a new project at every new stage of your growth, saving you time and money.&lt;/p&gt; &lt;/div&gt; &lt;p&gt;Another important aspect of any ETL workflow implementation is data quality and error management. We have an article explaining in detail &lt;a href=&quot;/blog/how-to-maintain-data-quality/&quot;&gt;how to guarantee the quality of the data&lt;/a&gt; loaded in your databases, and how to deal with validation errors. So, we won’t go into too much detail in this article, but here’s a list of principles to consider during implementation:&lt;/p&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/12.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 5. Error Detection and Data Validation&lt;/h4&gt; &lt;p&gt;During validation, the data being extracted is checked according to a predefined profile. This profile represents what a “good” data is for your project. The goal is to control data as early as possible to limit the computing time and avoid processing data that will be rejected later on. It also makes recovery easier as errors are detected early in the process.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/3.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 6. Recovery and Restartability&lt;/h4&gt; &lt;p&gt;Recovery and restartability address the capability of the workflow to resume after an error. It includes the process by which the data stays in a stable state following an error. That requires the use of database backup, as well as commit and rollback features. When we commit, we make permanent a set of changes in the code. Rollback, on the other hand, is the capability to return a program back to an earlier version. They are both used to manage workflow and errors, in combination with another practice known as idempotence.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/20.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 7. Idempotence&lt;/h4&gt; &lt;p&gt;An operation is defined as idempotent when it gives the same result after being called once or multiple times. In real life, the best example would be the elevator button: you have the same result whether you push it once or fifteen times. How can idempotence be relevant in data transformation, knowing how your data is always changing? Because sometimes, it doesn’t. If a data suddenly stops being updated, you still get the same results in your tables. The same would apply if your own transformation deployments were to stop. With idempotent transformations, you avoid a system failure when the ETL process itself fails. &lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/15.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 8. Data Lineage&lt;/h4&gt; &lt;p&gt;Data lineage helps identify which ETL steps a specific data point went through and to understand where it originated from, when it was loaded, and how it was transformed. It eases debugging and increase trust in the data by making the process transparent, thus validating the integrity of the end results. Thanks to lineage, we can guarantee the integrity of the data and the process that extracted and loaded it into your database.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/5.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 9. Auditing&lt;/h4&gt; &lt;p&gt;Checking your logs for potential mistakes is not enough to ensure that your load was a success. Your system should be designed to check for errors and to support auditing of your primary metrics (like the number of rows processed).&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/19.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 10. Script Configuration&lt;/h4&gt; &lt;p&gt;When executed, the configuration variable modifies a workflow behaviour. Variables are stored separately from the job so you can modify them without editing and deploying the scripts. Identifying the right variables is an integral part of any professional ETL job. For example, configuration variables contains parameters such as the server name and credentials, which should never be hardcoded in a job&lt;/p&gt; &lt;/div&gt; &lt;h3 id=&quot;four-best-practices-for-etl-workflow-to-schedule-monitor-and-maintain&quot;&gt;FOUR BEST PRACTICES FOR ETL WORKFLOW TO SCHEDULE, MONITOR, AND MAINTAIN&lt;/h3&gt; &lt;p&gt;Every data project includes an administrative component. We already covered &lt;a href=&quot;blog/schedule-maintain-web-scraper/&quot;&gt;the scheduling, monitoring, and maintaining of web scraping&lt;/a&gt;. However, these three important tasks also need to be executed at the ETL process level. Here’s a list of principles that will help your developer manage the workflows more efficiently. Most ETL software comes with a server edition that provides those four features. If you are looking for a technology agnostic (or cheaper) solution, &lt;a href=&quot;/contact/&quot;&gt;contact us&lt;/a&gt; for more details.&lt;/p&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/11.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 11. Orchestration&lt;/h4&gt; &lt;p&gt;Your developer will ensure that all the moving parts of your workflow come together to deal with the different nature, frequency, and cadence of your source data. It includes executing the different ETL modules and their dependencies, in the right order, along with logging, scheduling, alert monitoring, and managing code and data storage. This orchestration demands a high level of know-how, but also access to the right resources. You should never hesitate to ask for the services of an expert like us to help you implement your project.&lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/8.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 12. Metadata management&lt;/h4&gt; &lt;p&gt;Metadata is basically data about your data. They hold all kinds of information describing our ETL workflow. It can include information on where the data comes from, how many data points it contains and the data extraction strategy. Most importantly, a well-designed metadata system will maintain various versions of execution, including the status, the extraction and transformation methods used, the changes in source systems, etc. Thanks to the metadata, your developer can keep track of all these changes over several months or even years and your team will have all it needs to analyze the system more efficiently. &lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/18.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 13. Logging&lt;/h4&gt; &lt;p&gt;Every step of an ETL project must be logged using a central logging component. Relevant events are then recorded whether they happen before, during, or after extraction, transformation, or loading. &lt;/p&gt; &lt;/div&gt; &lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/14rules/1.jpg&quot; width=&quot;10%&quot; /&gt; &lt;h4&gt; 14. Code Management and Storage&lt;/h4&gt; &lt;p&gt;You should always keep track of your code and store it somewhere safe. There are three main reasons for this. First, you should always keep track of your code versions, so you can come back and restore your scripts after a bug or an error. Second, code management and storage will enable collaboration between members of your team and other external collaborators. And third, because your code needs to be kept separate from the execution environment. &lt;/p&gt; &lt;/div&gt; &lt;h3 id=&quot;and-with-that&quot;&gt;And with that&lt;/h3&gt; &lt;p&gt;With those fourteen best practices, you have everything you need to make sure your ETL workflow fits your organization’s needs. Whether your plan is to grow, to discover new markets, to know more about your competitors, to develop a business plan, or to introduce a new product to the market, data should always be at the cornerstone of your analysis. Selecting the best ETL tool will only be the first step to ensure your end data is reliable and relevant to your goals. Implementing and maintaining these tools and the process that binds them is paramount to your project success.&lt;/p&gt; &lt;p&gt;This article only scratches the surface of ETL design principles and best practices. Your developer will need to know which ones need to be applied, when they should be implemented, and at what range. Your developer needs to balance the robustness of the data pipeline and its development cost. But these principles and guidelines implemented at the right moment with the right goal in mind will guarantee the quality of your data, but will also help you manage an already complex process with more ease and fewer headaches.&lt;/p&gt; &lt;p&gt;At RefinePro, we have been helping clients implement ETL projects for years, and doing so, have developed a deep understanding of its many internal mechanisms. And we rely on such best practices to guarantee that our ETL workflows answer all our client’s needs.&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="ETL"/><category term="Strategy"/><summary type="html">Extracting, transforming, and loading (ETL) data is a complex process at the center of most organizations’ data extraction projects. As we saw in our article on web scraping and ETL, the implementation of an ETL workflow is a process that requires a lot of in-depth knowledge in several subfields of statistics and programming.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/1410092930.jpg"/></entry><entry><title type="html">How to Maintain Data Quality at Every Step of Your Pipeline</title><link href="https://refinepro.com/blog/how-to-maintain-data-quality/" rel="alternate" type="text/html" title="How to Maintain Data Quality at Every Step of Your Pipeline"/><published>2020-04-05T00:00:00-04:00</published><updated>2020-04-05T00:00:00-04:00</updated><id>https://refinepro.com/blog/how-to-maintain-data-quality</id><content type="html" xml:base="https://refinepro.com/blog/how-to-maintain-data-quality/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1410092930.jpg&quot; width=&quot;100%&quot; /&gt; &lt;p&gt;Maintaining the quality of your data is paramount to any web scraping or data integration project. Think about it: there’s absolutely no point in collecting a massive amount of data if you can’t rely on it to make sound decisions! And the only way to maintain high quality is by implementing quality checks and validation at every step of your data pipeline. As the saying goes: garbage in, garbage out!&lt;/p&gt; &lt;/div&gt; &lt;p&gt;We’ve discussed already what a good ETL (Extract Transform Load) tool is and what it should do, and we’ll now learn how data quality insurance fits into that process. ETL is the process of extracting, transforming and loading the data. It defines what, when, and how data gets from your source sites to your readable database. Data quality, on the other hand, relies on the implementation of a system from the early stage of extraction all the way to the final loading of your data into readable databases.&lt;/p&gt; &lt;p&gt;&lt;a href=&quot;blog/who-why-why-of-web-scraping/&quot;&gt;Choosing the right scraper&lt;/a&gt; and the right ETL tools will help you streamline this process, but these tools don’t automatically guarantee the quality of your end results. That’s why you need to work with a partner, like us, who will put in place all the proper checkpoints.&lt;/p&gt; &lt;h3 id=&quot;data-quality-in-your-etl-process&quot;&gt;DATA QUALITY IN YOUR ETL PROCESS&lt;/h3&gt; &lt;h4 id=&quot;extract&quot;&gt;Extract&lt;/h4&gt; &lt;p&gt;When you sat down to define &lt;a href=&quot;/expertise/web-scraping/&quot;&gt;your web scraping project&lt;/a&gt;, you made a list of sources you would be collecting the data from. Already, the choices you made will have an impact on the quality of the data. It’s important to always rely on trustworthy source sites that are relevant to your goals.&lt;/p&gt; &lt;p&gt;Don’t forget, &lt;a href=&quot;/blog/schedule-maintain-web-scraper/&quot;&gt;scheduling, maintaining, and monitoring&lt;/a&gt; are essential aspects to ensure your data is current. For example, if you always extract your competitor prices the day they put everything on sale, your data won’t reflect the real sale prices.&lt;/p&gt; &lt;p&gt;At the extracting phase, you know what your data is, and you should already implement scripts that will check its quality. This way, you allow the system to troubleshoot closer to the source itself, and you can take action immediately before your data are transformed.&lt;/p&gt; &lt;h4 id=&quot;transform&quot;&gt;Transform&lt;/h4&gt; &lt;p&gt;Transformation is when most of the quality checks are done. No matter what tool is used, it should at least perform the following tasks:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;&lt;strong&gt;Data profiling&lt;/strong&gt;: the data is analyzed in terms of quality, but also format, volume, etc.&lt;/li&gt; &lt;li&gt;&lt;strong&gt;Data matching and cleansing&lt;/strong&gt;: related entries are merged, and duplicates are eliminated.&lt;/li&gt; &lt;li&gt;&lt;strong&gt;Data enrichment&lt;/strong&gt;: the value of your data is extended by adding other relevant data to it.&lt;/li&gt; &lt;li&gt;&lt;strong&gt;Data normalization and validation&lt;/strong&gt;: the integrity of the data is checked, and validation errors are managed.&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;That last point is particularly important. Data validation is the process used to ensure the high quality of the data after it’s been cleaned and enriched. During the validation process, data is being checked against predefined rules, constraints, or routines regarding its meaningfulness and correctness. And to implement data security solution, dynamic profiling is, without a doubt, the best solution. Integrated into the validation process, it automatically profiles your dataset, including value type, mandatory field, number of records, and use them to classify data according to a baseline of what is “acceptable”: the profile. Once deployed, it automatically warns you when a data does not match the defined profile, without the need for manual handling.&lt;/p&gt; &lt;p&gt;When we sit down to create a project, we will help you define and document what a good profile is for your data. While developing the ETL workflow, we make sure the system continuously validates the results. And we implement a test-driven development process to &lt;strong&gt;ensure the data quality is not over-engineered.&lt;/strong&gt;&lt;/p&gt; &lt;h4 id=&quot;load&quot;&gt;Load&lt;/h4&gt; &lt;p&gt;At this point, you know your data. It’s been transformed to fit your needs and, if your quality check system is efficient, the data that reaches you is reliable. This way, you avoid overloading your database or data warehouse with unreliable or bad quality data, and you ensure that the end results have been validated. Your team can then use it to define strategies and make plans.&lt;/p&gt; &lt;h3 id=&quot;three-data-quality-strategies&quot;&gt;THREE DATA QUALITY STRATEGIES&lt;/h3&gt; &lt;p&gt;With validation and dynamic profiling, you ensure that the database you send to your team contains only data that’s usable and meets your most specific requirements. But in order to do that, you also need to deal with error validation. These errors occur when a data does not respect the validation rules established for the project. In other words, the data you extracted can’t get in the transformation step, or it doesn’t meet your requirements. This error can originate from either a bug in your workflow (you had the right data, but implementation did not work out) or changes in the data coming from the source (the extracted data structure did not match what was expected).&lt;/p&gt; &lt;p&gt;Depending on the on how critical is the data, we apply one of the following:&lt;/p&gt; &lt;ol&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Continue processing:&lt;/strong&gt; When a validation error occurs, processing continues with the incorrect records. The system logs the error to be reviewed later. This strategy works well for non-critical systems with a limited budget for data quality supervision.&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Partial processing:&lt;/strong&gt; When a validation error occurs, the processing continues and the erroneous record is rejected. The rejected records need to be reviewed and corrected before they can be reintroduced into the workflow. It is the most common approach. It protects downstream systems while ensuring good data keep flowing.&lt;/p&gt; &lt;/li&gt; &lt;li&gt; &lt;p&gt;&lt;strong&gt;Circuit breakers:&lt;/strong&gt; When an error occurs, the circuit opens, preventing low-quality data from propagating to downstream processes. Latest changes are rolled back, so the target system is closer to the version that existed before the execution of the ETL flow. In this case, the resulting data will not include all records, but records are guaranteed to be correct if present. This is the most conservative approach. However, you have to keep in mind that the system is on pause until someone addresses the error and restarts the system.&lt;/p&gt; &lt;/li&gt; &lt;/ol&gt; &lt;h3 id=&quot;why-data-quality-should-be-decoupled&quot;&gt;WHY DATA QUALITY SHOULD BE DECOUPLED&lt;/h3&gt; &lt;p&gt;We saw how your quality check system is integrated into the ETL process of a data pipeline. But in real life, your data quality insurance system should cover all the data pipeline. When we work with clients, we create scripts that support multiple data collection and transformation scripts (and so, numerous ETL processes).&lt;/p&gt; &lt;p&gt;How would that work? Let’s say you’re trying to gather data from your suppliers’ catalog to populate your eCommerce site. You could create ten different collections and transformation scripts for ten different suppliers, but only write a single data quality script to ensure quality.&lt;/p&gt; &lt;p&gt;The other good reason to implement ETL and quality checks as distinct processes is to ensure easy restartability. Collection and transformation can take hours to process. By having reliable data quality check different from the web scraping or ETL, you don’t have to restart the entire pipeline when you make a small correction, or you fix a bug.&lt;/p&gt; &lt;p&gt;If you rely on your processing steps (scraping or ETL) to check the quality of your data, you run the risk of having a system that is too cumbersome to manage, and so you run the risk of letting bad data slip through the cracks.&lt;/p&gt; &lt;h3 id=&quot;and-with-that&quot;&gt;And with that&lt;/h3&gt; &lt;p&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1117687502.jpg&quot; width=&quot;45%&quot; /&gt; Having a robust ETL tool supported by a great scraper is crucial to any data aggregation project. But to ensure that the end results meet your needs, you also need to make sure you have a &lt;a href=&quot;/expertise/data-quality-cleansing/&quot;&gt;quality check system in place&lt;/a&gt;.&lt;/p&gt; &lt;p&gt;When we sit down with you to define what your needs are, we make sure that the system generates data that is reliable, usable, but also relevant to your goals. We make sure that our scripts allow efficient error management in a timely manner. Because errors will occur: websites get updated all the time without notice, and unless your system helps you deal with these changes, you run the risk of having missing or erroneous data embedded in your datasets. Or you run the risk of your system crashing without knowing how to fix it. That’s why we &lt;a href=&quot;/expertise/migration-integration/&quot;&gt;develop data pipelines&lt;/a&gt; that integrate all the required validation processes to make sure the end results are what you’re expecting.&lt;/p&gt; &lt;p&gt;Also, we recommend that you take a look at a &lt;a href=&quot;https://www.toptal.com/database/data-warehouse-data-quality-process&quot;&gt;Toptal article&lt;/a&gt; on data quality in data warehousing and this &lt;a href=&quot;https://dataladder.com/benefits-data-matching/&quot;&gt;article on the benefit of data matching&lt;/a&gt; by Data Ladder&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="Data Quality"/><category term="Strategy"/><summary type="html">Maintaining the quality of your data is paramount to any web scraping or data integration project. Think about it: there’s absolutely no point in collecting a massive amount of data if you can’t rely on it to make sound decisions! And the only way to maintain high quality is by implementing quality checks and validation at every step of your data pipeline. As the saying goes: garbage in, garbage out!</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/1410092930.jpg"/></entry><entry><title type="html">The Who, the What, and the “With What” of Web Scraping</title><link href="https://refinepro.com/blog/who-why-why-of-web-scraping/" rel="alternate" type="text/html" title="The Who, the What, and the “With What” of Web Scraping"/><published>2020-03-12T00:00:00-04:00</published><updated>2020-03-12T00:00:00-04:00</updated><id>https://refinepro.com/blog/Who-Why-What-of-Web-Scraping</id><content type="html" xml:base="https://refinepro.com/blog/who-why-why-of-web-scraping/">&lt;div style=&quot;clear:both; min-height:250px&quot;&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1285494007.jpg&quot; width=&quot;100%&quot; /&gt; &lt;p&gt;Data is the new differentiator. It’s what you, a product owner, a marketing strategist, your local journalist, and a multimillionaire who already owns twelve successful companies all need. And web scraping is one way to get that data. &lt;/p&gt; &lt;/div&gt; &lt;p&gt;But where to start? Sure, the Internet will give you everything you need to know. Soon, you’ll come across lists of the best tools available, each with a name that will never make the list for best marketing decision of the year: Octoparse, Scrapy, BeautifulSoup, ParseHub, Mozenda … But how to choose? Looking at the description and the ratings is a good system when you’re buying shoes, but not when you’re trying to find the best scraper for your project.&lt;/p&gt; &lt;p&gt;Think about it: if you’re about to send something out there on the web to gather the reliable data you need, you have to make sure that the tool you’re using is the best for your project and your specific goals. A pair of flip flops, no matter how good the ratings, won’t get you far if you’re visiting Norway in December. And that’s why we put so much energy in helping our clients find the right tools for their projects. Over the years, we’ve tested many of them and you can now benefit from our experience. Here’s what we know.&lt;/p&gt; &lt;h3 id=&quot;the-crawler-crawling-and-the-scraper-scraping&quot;&gt;The crawler crawling and the scraper scraping&lt;/h3&gt; &lt;h4 id=&quot;web-scraping-101&quot;&gt;Web Scraping 101&lt;/h4&gt; &lt;p&gt;Web scraping is the process of fetching and extracting data from websites and downloading it in a usable format. It’s also referred to as “web harvesting,” “crawling,” “spidering,” and “web data extraction.” The process usually involves a &lt;strong&gt;scraper&lt;/strong&gt;, the tool designed to extract the data for the webpages, and the &lt;strong&gt;crawler&lt;/strong&gt;, or the spider, whose job is to browse the Internet to index and search for relevant content. Web scraping saves you the trouble of manually searching, downloading, and copying the data you need, and will work regardless of format. And by gathering big sets of data, you can help your organization grow by creating new products and innovating faster.&lt;/p&gt; &lt;h4 id=&quot;what-is-web-web-scraping-for&quot;&gt;What is Web Web Scraping For?&lt;/h4&gt; &lt;p&gt;Web scraping is used across an extensive range of industries, including risk management, retail, finance, sales and marketing, insurance, artificial intelligence, and journalism, to name only a few.&lt;/p&gt; &lt;p&gt;Web scraping can help you:&lt;/p&gt; &lt;ul&gt; &lt;li&gt;Gather data on your competitors and their products&lt;/li&gt; &lt;li&gt;Optimize your customer relationship management&lt;/li&gt; &lt;li&gt;Detect fraud&lt;/li&gt; &lt;li&gt;Feed a natural language system&lt;/li&gt; &lt;li&gt;Train machine learning models&lt;/li&gt; &lt;li&gt;Generate more and better leads&lt;/li&gt; &lt;li&gt;Perform reliable competitive analysis&lt;/li&gt; &lt;li&gt;Monitor your reputation&lt;/li&gt; &lt;li&gt;And so much more&lt;/li&gt; &lt;/ul&gt; &lt;p&gt;Web scraping has opened the door to big data in a world where every market research and business strategy relies on it. Think about it: all your strategies, plans, and insights into the future rely on data. Web scraping offers the privilege of non-discrimination.&lt;/p&gt; &lt;blockquote&gt; &lt;p&gt;As long as you adhere to the standards in place and use the right system with enough data warehousing capacity, there are no limits to the amount of data you can collect.&lt;/p&gt; &lt;/blockquote&gt; &lt;p&gt;But whether you’re trying to figure out how your consumers feel about your product or trying to gather data from hundreds of websites in real-time, you’ll need different tools and capacities. The first step of every web scraping project is to decide whether you’d like to take care of everything internally or with a partner, like us. And that’s only the first question of a long list of things you need to figure out before you decide on a scraper.&lt;/p&gt; &lt;h3 id=&quot;our-10-criteria-to-evaluate-a-web-scraping-software&quot;&gt;Our 10 criteria to evaluate a Web Scraping software&lt;/h3&gt; &lt;p&gt;When the time comes to choose the right scraper, you should base your decision on specific criteria, and not on benchmarks. These criteria are largely defined by your project and capacities, and by your choice of externalizing or keeping everything in-house. To make the comparison easier we prepared a &lt;a href=&quot;/blog/download-web-scraping-comparison-form/&quot;&gt;ready to fill PDF form to evaluate three web scraping solutions&lt;/a&gt;.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;1. The solution maturity&lt;/strong&gt;. You want to make sure you invest in a technology that is going to be there in the long run and actively supported. Check when the software was initially released and how often it’s been updated. Make sure the documentation and custom support is available. Do you want an open source or proprietary solution?&lt;/p&gt; &lt;p&gt;&lt;strong&gt;2. The development environment&lt;/strong&gt;. For example, the choice of Windows, Mac or Linux and browser-based for SaaS. (SaaS are “Software as a service”, a software model licensed on a subscription basis and hosted by a third-party provider.) Take also into account how complex is the software and what skills do you need to write a scraper? How fast can you ramp-up new employee?&lt;/p&gt; &lt;p&gt;&lt;strong&gt;3. The execution platform and hosting options.&lt;/strong&gt; Once the project developed where can you execute it? Either public, private or hybrid clouds. The main question here is: do you want to rely on a third-party infrastructure to collect critical data or do you need to keep everything in-house?&lt;/p&gt; &lt;p&gt;&lt;strong&gt;4. The possibility to circumvent CAPTCHA&lt;/strong&gt;. That can be done using a third-party CAPTCHA-solving service.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;5. The ability to fine-tune proxy management and rotation&lt;/strong&gt;. This could, for example, help you select the countries from which requests will come and support for residential IP addresses. Good web scraper lets you connect with third-party proxy providers.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;6. The ability to handle advanced anti-scraping features&lt;/strong&gt; including device fingerprint anonymization and fine-tune browser/profile management.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;7. The capacity to add custom scripts.&lt;/strong&gt; By adding new pieces of code, we can extend the software capabilities.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;8. API and Workflow integration&lt;/strong&gt;. How easily can you integrate your web scraping project into your workflow? The availability of an API or external connectors to configure and project but also retrieve the data.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;9. Scheduling, monitoring, and maintaining.&lt;/strong&gt; These are crucial to any web scraping project, so your tool should allow you to perform all three according to your needs.&lt;/p&gt; &lt;p&gt;&lt;strong&gt;10. Pricing.&lt;/strong&gt; At the end of the day, we all have a budget to respect.&lt;/p&gt; &lt;p&gt;So, ultimately, it’s your project, and most importantly, its schedule, scale, and eventually budget, that will determine what tools you should use.&lt;/p&gt; &lt;div style=&quot;text-align:center;&quot;&gt; &lt;a href=&quot;/blog/download-web-scraping-comparison-form/&quot; class=&quot;button special&quot;&gt;Download our editable form for your personal use&lt;/a&gt;&lt;/div&gt; &lt;p&gt;&lt;br /&gt;&lt;/p&gt; &lt;h3 id=&quot;web-scrapers-our-favorites&quot;&gt;Web Scrapers: Our Favorites&lt;/h3&gt; &lt;p&gt;Once you know more precisely the kind of tool you’re looking for, it’s time to start shopping. You’ll find a plethora of websites listing all the best tools with their pros and cons. The truth is, however, that few of them have tested those tools as much as we have, with projects that differ in goals, scale, and complexity. So, to make it easier for you, we’ve gathered a list of our three favorites, with a short description and comparative chart. If you’d like to know more, or you’re not sure how to decide, &lt;a href=&quot;/contact/&quot;&gt;contact us!&lt;/a&gt;&lt;/p&gt; &lt;h4 id=&quot;parsehub&quot;&gt;ParseHub&lt;/h4&gt; &lt;p&gt;&lt;br /&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/toolbox/parsehub.png&quot; /&gt; &lt;a href=&quot;/toolbox/parsehub/&quot;&gt;ParseHub&lt;/a&gt; is a point and click web scraping software that manages all projects on its infrastructure. This system means that you’re completely dependent on their infrastructure, yet you benefit from not having to make any of the usual set-up investment to provision environments. Their plan includes the maintenance of web scraping servers and proxy networks, preventing unexpected costs as you scale.&lt;/p&gt; &lt;p&gt;ParseHub is a bit more expensive than its counterparts (it does offer a free version for small and simple projects, though), but it’s a great option for easy projects with high volumes of data.&lt;/p&gt; &lt;h4 id=&quot;content-grabber&quot;&gt;Content Grabber&lt;/h4&gt; &lt;p&gt;&lt;br /&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/toolbox/content_grabber.png&quot; width=&quot;33%&quot; /&gt; &lt;a href=&quot;/toolbox/content_grabber/&quot;&gt;Content Grabber&lt;/a&gt;, a point and click web scraping software developed by Sequentum. It provides a robust and scalable solution to collect data from complex websites and offers the advantage of being deployed on-premises, on Windows servers. We can help you manage your project without relying on a third-party vendor. And by having full control over the infrastructure, we can meet the most restrictive data privacy and security requirements.&lt;/p&gt; &lt;h4 id=&quot;puppeteer&quot;&gt;Puppeteer&lt;/h4&gt; &lt;p&gt;&lt;br /&gt; &lt;img class=&quot;alignleft&quot; src=&quot;/images/toolbox/puppeteer.png&quot; /&gt; &lt;a href=&quot;https://pptr.dev&quot;&gt;Puppeteer&lt;/a&gt; is a headless browser that uses DevTools Protocol to communicate with Chrome or Chromium. We call it “headless” because it was designed to be used by machines, not humans. It has no user interface and its main goal is to allow programs to read and interact with it. Like Content Grabber and ParseHub, Puppeteer is well designed for large projects, but it offers a complete solution control for complex websites using advanced anti-scraping features.&lt;/p&gt; &lt;p&gt;But complex features also mean a more complex design, so Puppeteer is not for the neophytes. It requires a high level of expertise and you’ll need a trained developer to create the project and maintain it. But as we’ve mentioned many times before, it’s a mistake to think of &lt;a href=&quot;blog/schedule-maintain-web-scraper/&quot;&gt;web scraping as a product: it’s a service&lt;/a&gt;. So, if your project requires complex extraction, make sure to ask an expert like us to develop and schedule your project, but also to maintain and monitor it.&lt;/p&gt; &lt;h4 id=&quot;the-comparison&quot;&gt;The Comparison&lt;/h4&gt; &lt;p&gt;Here’s a slightly more complete table that compares the three tools described above.&lt;/p&gt; &lt;table&gt; &lt;tbody&gt; &lt;tr&gt; &lt;td&gt; &lt;/td&gt; &lt;td&gt;&lt;strong&gt;&lt;a href=&quot;/toolbox/parsehub/&quot;&gt;ParseHub&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;&lt;strong&gt;&lt;a href=&quot;/toolbox/content_grabber/&quot;&gt;Content Grabber&lt;/a&gt;&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;&lt;strong&gt;Puppeteer&lt;/strong&gt;&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;1. Solution Maturity&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Mature proprietary software- Started in 2015. New version released in Fall 2019.&lt;/td&gt; &lt;td&gt;Mature proprietary software - Released 2015. Follow Visual Web Ripper released in early 2000.&lt;/td&gt; &lt;td&gt;Mature open source - Started in 2017. Large community support with 250+ developers&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;2. Development Environment&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Easy Point and click software for Mac, Windows, and Linux&lt;/td&gt; &lt;td&gt;Easy Point and click software for Windows only&lt;/td&gt; &lt;td&gt;Complex, using a developer envirnoment on Mac, Windows, and Linux&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;3. Excution Platform &amp;amp; Hosting&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;ParseHub Platform (SaaS)&lt;/td&gt; &lt;td&gt;Self-hosted Windows server&lt;/td&gt; &lt;td&gt;Self-hosted Linux server&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;4. Captcha&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Basic resolution included, possible to connect to third-party&lt;/td&gt; &lt;td&gt;Connect with third-party&lt;/td&gt; &lt;td&gt;Connect with third-party&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;5. Proxy&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Basic proxy included, possible to connect to a third-party provider&lt;/td&gt; &lt;td&gt;Connect with third-party&lt;/td&gt; &lt;td&gt;Connect with third-party&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;6. Anti Scraping&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Limited&lt;/td&gt; &lt;td&gt;Advanced&lt;/td&gt; &lt;td&gt;Advanced&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;7. Custom Scripts&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Javascript and regular expression to select content only&lt;/td&gt; &lt;td&gt;Extends the software with C#, regular expression, VB, Python 3&lt;/td&gt; &lt;td&gt;JavaScript environment, but everything is code-based!&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;8. API and Workflow Integration&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Yes—to orchestrate (start, stop, and pass parameter) and retrieve data&lt;/td&gt; &lt;td&gt;Yes—to orchestrate (start, stop, and pass parameters). &lt;br /&gt; &lt;br /&gt; Support data export in multiple formats&lt;/td&gt; &lt;td&gt;No—you need to orchestrate the script and manage your data export yourself&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;9. Ease to schedule and monitor&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Easy - everything is supported by ParseHub&lt;/td&gt; &lt;td&gt;Medium - You need to provide the infrastructure, then everything is managed via Content Grabber Agent Control Center&lt;/td&gt; &lt;td&gt;Not available. You need to provide&lt;/td&gt; &lt;/tr&gt; &lt;tr&gt; &lt;td&gt;&lt;strong&gt;10. Pricing&lt;/strong&gt;&lt;/td&gt; &lt;td&gt;Self Service. &lt;a href=&quot;https://parsehub.com/pricing&quot;&gt;Prices are available online&lt;/a&gt;.&lt;/td&gt; &lt;td&gt;Enterprise Solution. Contact Sequentum for a quote.&lt;/td&gt; &lt;td&gt;Free and Open Source.&lt;/td&gt; &lt;/tr&gt; &lt;/tbody&gt; &lt;/table&gt; &lt;div style=&quot;text-align:center;&quot;&gt; &lt;a href=&quot;/blog/download-web-scraping-comparison-form/&quot; class=&quot;button special&quot;&gt;Download our form to make your own analysis&lt;/a&gt;&lt;/div&gt; &lt;p&gt;&lt;br /&gt;&lt;/p&gt; &lt;h3 id=&quot;and-with-that&quot;&gt;And With That&lt;/h3&gt; &lt;p&gt;&lt;img class=&quot;alignleft&quot; src=&quot;/images/blog/1409055989.jpg&quot; width=&quot;45%&quot; /&gt; Web scraping can be easy. But to be easy, it necessitates a plan and a vast amount of technical know-how. Even when trying to choose the right tool, you should always seek the advice of a professional. Our goal here at RefinePro is to make sure your web scraping project will help you meet your long-term needs and that it will scale with you.&lt;/p&gt; &lt;p&gt;It’s easy to make a bad choice with web scraping and to send your team on a wild goose chase. Using the wrong tool to gather the wrong kinds of data on the wrong website, using wrong techniques that get you kick out are costly mistakes.&lt;/p&gt; &lt;p&gt;So instead of chasing a wild goose, give us a &lt;a href=&quot;/contact/&quot;&gt;call or email us&lt;/a&gt; to tell us about your project. And if you’re still not convinced you need help, go read our article about &lt;a href=&quot;/blog/schedule-maintain-web-scraper/&quot;&gt;the medium- and long-term implications of web scraping&lt;/a&gt;. You’ll learn how the launching of a web scraping project is only the beginning of a long story that can only end well if it involves ongoing and constant scheduling, monitoring, and maintenance.&lt;/p&gt; &lt;p&gt;We specialize in data. And we can help you make the best of it. Do you have any questions? &lt;a href=&quot;/contact/&quot;&gt;Contact us!&lt;/a&gt;&lt;/p&gt; &lt;div&gt; &lt;section class=&quot;special&quot;&gt; &lt;p&gt; &lt;!-- Contact CTA. Self-sufficient by design: it must render identically at all 23 call sites regardless of what wraps it, because the wrappers are inconsistent and two of them are invalid HTML. Before this was made self-sufficient there were three different renderings: - 9 evergreen posts &lt;div&gt;&lt;section class=&quot;special&quot;&gt;&lt;p&gt; centered, +2em above - toolbox / expertise / offering &lt;div&gt;&lt;section class=&quot;special&quot;&gt; centered, no space - index, platform, download posts bare include LEFT-aligned, no space The two mechanics behind that spread: - `section.special { text-align: center }` (css/main.scss:248). `.rp-cta--text` is width:50% inline-block with its own text-align:left, so the text/button pair has slack and the inherited alignment is visible. - `p { margin: 0 0 2em 0 }` (css/main.scss:118, element-margin: 2em). A &lt;section&gt; inside a &lt;p&gt; forces the parser to close the paragraph, leaving an EMPTY &lt;p&gt; whose bottom margin became the space above the CTA. That invalid markup was load-bearing, not inert. Hence the two inline declarations below: they replace what the wrappers were accidentally supplying. Inline rather than Sass, to avoid recompiling the whole stylesheet. `.wrapper.rp-cta` already contributes `padding: 3em 0` (_sass/rp-style.scss:186). --&gt; &lt;section id=&quot;contact-cta&quot; class=&quot;rp-cta wrapper style1&quot; style=&quot;text-align:center; margin-top:2em;&quot;&gt; &lt;div class=&quot;inner&quot;&gt; &lt;div class=&quot;rp-cta--text&quot;&gt;Got a project or idea in mind? &lt;br /&gt;We have the experts to make it happen. &lt;/div&gt; &lt;a href=&quot;/contact/&quot; class=&quot;rp-cta--button button special&quot;&gt;Tell Us About It&lt;/a&gt; &lt;/div&gt; &lt;/section&gt; &lt;/p&gt; &lt;/section&gt; &lt;/div&gt;</content><author><name>martin</name></author><category term="Web Scraping"/><summary type="html">Data is the new differentiator. It’s what you, a product owner, a marketing strategist, your local journalist, and a multimillionaire who already owns twelve successful companies all need. And web scraping is one way to get that data.</summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://refinepro.com/images/blog/"/></entry></feed>