<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:cc="http://cyber.law.harvard.edu/rss/creativeCommonsRssModule.html">
    <channel>
        <title><![CDATA[Stories by Ievgen Redko on Medium]]></title>
        <description><![CDATA[Stories by Ievgen Redko on Medium]]></description>
        <link>https://medium.com/@ievred?source=rss-fde5e3dd8903------2</link>
        <image>
            <url>https://cdn-images-1.medium.com/fit/c/150/150/2*iQGwTuIBuEBKk-UjC9IZSQ.jpeg</url>
            <title>Stories by Ievgen Redko on Medium</title>
            <link>https://medium.com/@ievred?source=rss-fde5e3dd8903------2</link>
        </image>
        <generator>Medium</generator>
        <lastBuildDate>Mon, 03 Aug 2026 17:16:57 GMT</lastBuildDate>
        <atom:link href="https://medium.com/@ievred/feed" rel="self" type="application/rss+xml"/>
        <webMaster><![CDATA[yourfriends@medium.com]]></webMaster>
        <atom:link href="http://medium.superfeedr.com" rel="hub"/>
        <item>
            <title><![CDATA[Marvel at GPT-4 as at your own very self]]></title>
            <link>https://ievred.medium.com/marvel-at-gpt-4-as-at-your-own-very-self-4a099828f15b?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/4a099828f15b</guid>
            <category><![CDATA[artificial-intelligence]]></category>
            <category><![CDATA[human-intelligence]]></category>
            <category><![CDATA[machine-learning]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Wed, 22 Mar 2023 21:36:21 GMT</pubDate>
            <atom:updated>2023-03-22T21:36:21.763Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*2ipvXFVSmnZw4ynP" /><figcaption>Photo by <a href="https://unsplash.com/es/@jupp?utm_source=medium&amp;utm_medium=referral">Jonathan Kemper</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><p>When I approached my wife 85 years old grandfather to ask whether I could go to the bakery instead of him today, he replied with a firm “Surtout pas !” which in translation from French means “Don’t you even think about it!”. This kind of reply was not because he doesn’t like me (I know he does) or anything of that sort — it is just that he enjoyed those daily activities that constituted the fiber of his life at this stage: bike to the bakery every morning, get the bread, go to the pier to have a chat with local fishermen, bike back home and work in his endless garage until his 84 years old wife calls him to lunch.</p><p>Occasionally, we do crafts together, and he tells me stories about their village life. This time, he talks about a new crêperie that opened last year but was too unauthentic by the high Brittany standards and went bankrupt in less than eight months. He says, “I couldn’t understand why they opened it here in the first place. It was obvious nobody was going to dine there.” I often hear him saying, “I like to understand things,” when he talks about a random business affair, about blockchain or dismantling the old swing and replacing it with his brother’s one having different dimensions. In his worldview, the reasoning and mental exercise behind the process of getting from point A to point B is as crucial as getting things done.</p><p>We usually stay for a couple of days and then go home with batteries fully recharged for a new round of our busy Parisian life. At work, my colleagues wonder what features GPT-4 — a new AI-based large language model developed by OpenAI — will have. When it is finally released, my Twitter feed and work chat instantly fill with praising words about its revolutionary multi-modal inputs and reasoning capacity, scoring higher than an average contender applying to Stanford University. Objectively, I understand that it is a huge achievement, yet somehow I fail to feel inspired.</p><p>While walking home after work, I try to unravel this thought and this lack of excitement. Slowly, I realize that most of the revolutionary things that GPT-4 does are those that I enjoy doing as well. Many people use it to write code, yet I learn a lot from coding myself; some also ask it to write down their thoughts clearly and more fluently, but I immensely enjoy writing and struggling with those things too. Being of a non-entrepreneurial nature, I avoid the rabbit hole of dreaming about the myriads of its potential commercial applications. In the end, I abandon my internal search and convince myself that the beauty of this engineering feat for me is more of a symbolic nature: it should incite us to marvel at how narrow the gap between what we are and what we are capable of creating is. Satisfied with this thought, I go back to my daily routines and wonder no more.</p><p>Then, one day I stumbled upon a tweet mentioning that one of the open problems in the field of combinatorial mathematics was solved after remaining unanswered for almost a century. In the tweet thread of one famous British mathematician, I see a photo of six guys sitting on a wooden bench around a table and drinking a pint under the March grey sky, celebrating their incredible achievement. A question then strikes me: if we are to celebrate the intelligence of GPT-4 then how excited should we get over achievements like this? Or even over anyone around us who carries the pieces of intelligence that GPT-4 was trained on in the first place? Those who were willing to understand and give it those precious bits of sequential reasoning it builds upon now? It almost feels surreal in some sense: as if we were standing in front of the Mona Lisa painting in the Louvre but we’re looking at its photo on a tiny screen of our smartphones instead.</p><p>In a month or so, I may be coming back to the Atlantic coast, and my wife’s grandfather will once again forbid me to interfere with his daily routine. I’ll join him on the market day though, and we’ll talk about the usual small things on the way there, those that GPT-4 would be good at. I’ll mention it to him, and he’ll probably know about it from an article that he’d read in a newspaper. “Would you want it to write your emails for you?” I would ask. “Surtout pas !” he would reply, and we’ll go back home with the baskets full of delicious food. We’ll live the day and get through numerous mundane things filling the time we have to spare. We’ll laugh together, drink red wine, and, when the day will draw to its end, we will marvel with my wife at how lucky we are to live through all these precious little moments.</p><p>In the end, I agree that GPT-4 may be well worth marveling at. But as far as I’m concerned, so do we, all of us whom it strives to mimic so well.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=4a099828f15b" width="1" height="1" alt="">]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Reviewing for Machine Learning Conferences Explained]]></title>
            <link>https://medium.com/data-science/reviewing-for-machine-learning-conferences-explained-f73bc037babc?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/f73bc037babc</guid>
            <category><![CDATA[review]]></category>
            <category><![CDATA[conference]]></category>
            <category><![CDATA[research]]></category>
            <category><![CDATA[science]]></category>
            <category><![CDATA[machine-learning]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Mon, 07 Dec 2020 15:46:58 GMT</pubDate>
            <atom:updated>2020-12-07T17:31:36.673Z</atom:updated>
            <content:encoded><![CDATA[<h4>From reading a paper for the first time to writing its complete review in a single Medium article</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*CJgRxx7tSholPAMW" /><figcaption>Photo by <a href="https://unsplash.com/@markuswinkler?utm_source=medium&amp;utm_medium=referral">Markus Winkler</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><p>Peer-reviewing is the cornerstone of modern science, and almost all major conferences in machine learning (ML), such as NeurIPS and ICML, rely on it to decide on whether submitted papers are relevant to the community and original enough to be published there. Unfortunately, with the <a href="https://medium.com/@dcharrezt/neurips-2019-stats-c91346d31c8f">exponentially increasing number of submitted articles over the last ten years</a>, the reviewing quality has been dropping just as fast, with one-line reviews became widespread. You’ve probably been there already if you’ve ever submitted a paper to one of these conferences: after having worked very hard for months on what you felt like was a brilliant idea, you receive awful, useless, and (worse of it all) ironic reviews meaning that you will have to go through the submission process all over again without any hint on what was wrong with your paper in the first place.</p><p>Geoffrey Hinton, the famous Turing award winner for his contributions to the machine learning and AI fields, gave one of the reasons for why this happens in his interview to Wired journal in 2018:</p><blockquote>Now if you send in a paper that has a radically new idea, there’s no chance in hell it will get accepted, because it’s going to get some junior reviewer who doesn’t understand it. Or it’s going to get a senior reviewer who’s trying to review too many papers and doesn’t understand it first time round and assumes it must be nonsense. Anything that makes the brain hurt is not going to get accepted. And I think that’s really bad.</blockquote><p>While senior reviewers have little to no excuse for such behavior (why would you voluntarily agree to review if you do not have time to do it properly?!), junior reviewers may simply not know how to write a good thoughtful review. Conference organizers usually provide <a href="https://neurips.cc/Conferences/2020/PaperInformation/ReviewerGuidelines">helpful guidelines</a> with examples from reviews gathered over the years, but this cannot explain how to write a full review from scratch: starting from reading the submitted article for the first time and to finalizing your review and submitting it on the conference website. As I happen to have won several so-called “Top reviewer” awards previously (IJCAI’18, NeurIPS’19, ’20), I would like to explain below how I proceed when I review papers hoping that it will be useful for people who may need such guidance.</p><p>I wrote this article based on the Research Methodology course that I teach to Master’s degree students in machine learning. One of its lectures goes as follows: we read the paper together paragraph by paragraph and I explain to what particular parts of it a reviewer has to pay attention. As an example, I use the paper from the ICLR’19 conference entitled “<em>Learning what and where to attend with humans in the loop</em>” (its first submitted version can be found <a href="https://openreview.net/references/pdf?id=r17sOn9FX"><strong>here</strong></a>). I chose this paper for two reasons: 1) it is not in one of my areas of primary expertise, and 2) it remains largely accessible to anybody having a general background in ML. I thought that the first point was very important as most of the future Ph.D. students will start their reviewer’s career in similar conditions and without a longstanding prior experience in any particular ML field.</p><p>I now propose you to follow me through the paper in order to understand how to write a review for it. To do this, I suggest you read full sections of the paper indicated in the titles below before reading my comments on it.</p><h3><strong>Abstract</strong></h3><p>The abstract is one of the most important things in the paper for a reviewer as it gives a general outline of what he/she will find in it. When reading this part, I note every promise made by authors and expect that authors supported it by facts in the main body of their work. Let’s see the abstract of our paper.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*GgmVQ2KAZd_euh9gKpGXUA.png" /><figcaption>Image by Author based on the original <a href="https://openreview.net/forum?id=BJgLg3R9KQ">paper</a>.</figcaption></figure><p>I put important things in bold here. What information this abstract gives me as a reviewer? First, it defines the <strong>general area of the submission,</strong> which is the study of attention mechanisms in DCNs. Second, and most importantly, it puts forward two claims which I will want to verify, namely: 1) attention mechanisms with human supervision <strong>significantly improve</strong> the DCN’s performance, 2) learned features, in this case, are more <strong>interpretable</strong>. I note it and move on to the introduction.</p><h3>Introduction</h3><p>The introduction is an extended version of the abstract that includes <strong>hints to previous works</strong> and provides <strong>more details on the proposed contribution</strong>. In this paper, the introduction contains several things that attract my attention.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/970/1*PzQWm7rQMFRm89WksdFr3Q.png" /><figcaption>Images by Author based on the original <a href="https://openreview.net/forum?id=BJgLg3R9KQ">paper</a>.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/968/1*O-x8EcKrPx-5dX1EUznyfg.png" /><figcaption>Images by Author based on the original <a href="https://openreview.net/forum?id=BJgLg3R9KQ">paper</a>.</figcaption></figure><p>First, I identify several <strong>closely related prior works</strong> mentioned numerous times in the second paragraph, namely: (Linsley et al. ’17) and (Jiang et al. ‘15). As a reviewer, I would now briefly go through the contents of these two papers with a particular emphasis on the first one because 1) it is more recent and most likely will include a comparison to other related works mentioned in the introduction, and 2) the authors compare to it <strong>singularly</strong>.</p><p>Second, I note the <strong>positioning</strong> of proposed contributions <strong>w.r.t. state-of-the-art</strong>, namely: 1) the authors propose a more efficient strategy implemented on ClickMe.ai platform to obtain attention maps for large-scale datasets when compared to Salicon dataset and Linsley et al.’s work; 2) the authors propose a novel module for DCNs based on the idea of combining global contextual guidance with local saliency; 3) the authors improve performance with human-in-the-loop attention. Once again, as a reviewer, I will now seek arguments that support each of these claims.</p><h3>Section 2: Description of ClickMe.ai</h3><p>This section is very important, as it is entirely devoted to supporting the first claim mentioned above. On the one hand, it is supposed to show that the proposed strategy used to collect attention maps scales better than previous work. On the other hand, it should show that the obtained “top-down” maps are superior to “bottom-up” ones collected previously. Here is my resume for the first part.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/681/1*IwD8_FPiMbNqRs2g9txy1g.png" /><figcaption>Image by Author based on the original <a href="https://openaccess.thecvf.com/content_ICCV_2017_workshops/papers/w40/Linsley_What_Are_the_ICCV_2017_paper.pdf">paper</a>.</figcaption></figure><p>You may note that I put the ClickMe.ai strategy proposed by the authors as a <strong>strength of the paper</strong> as it involves only one human being contrary to two from Linsley et al., and allows to collect more attention maps. A downside to this is that the comparison with Linsley at al. allows me to <strong>discover the identities</strong> of the authors who mention ClickMe.ai in their previous paper.</p><p>Here is my summary of the second claim.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/611/1*fRK-_FqTPLyc5mm860Wvcw.png" /><figcaption>Source: original <a href="https://openreview.net/forum?id=BJgLg3R9KQ">paper</a>.</figcaption></figure><p>As shown above, “top-down” features (ClickMe maps) seem to perform better than “bottom-up” features (Salicon maps) when revealed to human observers. This supports the author’s claim about their superior performance w.r.t. the maps from the Salicon dataset. So far, I am only praising the strengths of the paper, but is there something to say about the weaknesses? Here are some of my remarks to be included in the review.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/751/1*CJRn7Heco1421Opni9bWMg.png" /><figcaption>Image by Author.</figcaption></figure><p>The authors say that the ClickMe game scales better than the Clicktionary game of Linsley et al., but they never mention how many maps were collected using the latter. The second point is that the authors also use those maps that didn’t allow DCN to recognize the object correctly. Is it reasonable? Why do we consider these maps useful later? The other two minor points are that the authors often talk about “top-down” and “bottom-up” features, but they never explain the difference between the two (I had to google it). Finally, the authors claim that these features are “sufficient for human object recognition”. This may be too strong of a statement as the overall recognition accuracy never reaches 70%, which is far from what is considered human-level performance. I note it and move to Section 3.</p><h3>Section 3: Proposed network architecture</h3><p>This section describes the module inspired by the idea of “combining local saliency and global contextual signals to guide attention towards image regions that are diagnostic for object recognition”. I am not an expert in attention mechanisms for DCNs, and I cannot judge the soundness of what is proposed by the authors and its novelty. At this point, I start to think that my <strong>confidence score</strong> for this paper won’t be very high if I were its official reviewer and that I would have to <strong>indicate</strong> it clearly in my review to the <strong>area chair</strong> (AC). Despite this, I still notice the following phrase of the authors:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/963/1*vPDWCAS3c74d1tmvTbsL7A.png" /><figcaption>Images by Author based on the original <a href="https://openreview.net/forum?id=BJgLg3R9KQ">paper</a>.</figcaption></figure><p>They do not explain how they choose these layers, while omitting low-level layers altogether: an ablative study might be useful to back it up in this context. One more thing to add to the review as such discussion can be highly useful for researchers who may decide to implement their module for architectures other than ResNet-50.</p><h3>Section 4: Training with humans-in-the-loop</h3><p>This section presents most of the experimental results for the architecture proposed in Section 3 with additional regularization that forces learned maps to look similar to those provided by human participants of ClickMe.ai. Here is my short resume for this part.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/694/1*Rn-9gdDUbeAIBrWbBq3mcA.png" /><figcaption>Images by Author.</figcaption></figure><p>I note that most of the results indeed seem to back up the third claim from authors: human-in-the-loop supervision improves the performance on popular object recognition datasets. Even though I would tend to be convinced by the experimental results, I still notice several inconsistencies.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/753/1*2WxR2gbKP170u9i5erE0yg.png" /><figcaption>Images by Author.</figcaption></figure><p>The <strong>first</strong> remark is rather obvious. Why using the magic number 6 for the regularization parameter? The <strong>second</strong> question is related to Table 1 from the paper that shows significant improvement in terms of both the classification accuracy on the ILSVRC12 dataset and the ability to learn features similar to ClickMe maps. What’s inconsistent here? Well, the latter improvement seems to be quite obvious to me as it merely indicates that the regularization forcing learned features to look like ClickMe maps works well. Other baselines do not particularly seek to force such behavior, and this performance gain should be presented rather as an argument justifying the chosen regularization strength. <strong>Third</strong>, the authors mention that with a reduced set of ClickMe maps (Table 4 in Appendix), their method also performs better than all other baselines, but one can see that in this case, the performance gap becomes very small. <strong>Finally</strong>, the authors mention that “Without additional training, the model’s attention localized foreground objects in Microsoft COCO 2014 (Lin et al., 2014)” but do not provide quantitative results on this dataset and show only 6 derived maps in Figure 4 (this was improved in the published version).</p><h3>Putting it all together</h3><p>After explaining how I went through this paper, it is now time to put it all into a review ready to be submitted to the conference website. As required by many conferences, I start with the summary of the paper.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/973/1*SqdXgyqbuMgubf9xIC0BbQ.png" /><figcaption>Image by Author.</figcaption></figure><p>Note that the summary is very important as it shows the authors that you understand their work. Then, I provide its strengths and weakness.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/931/1*Ely9p7UVBYbVWUO-FybdRA.png" /><figcaption>Image by Author.</figcaption></figure><p>I find it crucial to give some positive feedback even if I plan to suggest rejecting the paper in the end. This shows the authors what parts of their work were appreciated by the reviewers. I then continue with several detailed comments.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/968/1*Ej1p0h1XbWo3hKRmq9VPPQ.png" /><figcaption>Image by Author.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/968/1*tNOXMW8k5wrlO1judRIffQ.png" /><figcaption>Image by Author.</figcaption></figure><figure><img alt="" src="https://cdn-images-1.medium.com/max/952/1*F60-Fzj4paDWgVRaRY6dNA.png" /><figcaption>Image by Author.</figcaption></figure><p>You may note that all of the review’s contents are just the remarks that I was writing down while reading the paper. Usually, I will make a first draft of a review in around 3 hours and then I will go back to the paper at least twice before the deadline to make sure that I didn’t miss something.</p><h3>What do other reviewers say?</h3><p>The good thing about how the reviewing process works nowadays is that you can often see the reviews for a given paper once it was accepted/rejected. In the case of this submission, you can check the reviews<strong> </strong><a href="https://openreview.net/forum?id=BJgLg3R9KQ"><strong>here</strong></a><strong>.</strong> After reading them, you may notice that other reviewers raise similar concerns to what I mention in my review, namely: 1) lack of motivation/justification for many design choices and 2) qualitative results for the interpretability. Note also that similar to me Reviewer 2 admits that he/she is not an expert in attention mechanisms for DCNs and puts a confidence score of 3/5 to indicate this to the area chair. This is very important as an unknowledgeable reviewer with high confidence is a nightmare for both the authors and the ACs. And it goes the other way around too: if you review a paper from your narrow area of expertise, you should clearly indicate it so that the AC will be able to identify the most informative reviews.</p><h3>What do the authors do then?</h3><p>I did the review of the first submitted version of this paper on purpose so that you can see the <a href="https://openreview.net/pdf?id=BJgLg3R9KQ"><strong>camera-ready version</strong></a> submitted by the authors once their paper was accepted. You may notice several differences in it compared to the first version: the title has changed to “Learning what and where to attend” as suggested by Reviewer 1, and many details were added throughout the whole text to make the paper clearer following the reviewers’ remarks (you can see it in <a href="https://openreview.net/references/pdf?id=BJBY3__0Q"><strong>the diff file</strong></a> between the final version and the original version). Overall, it shows you that your duty as a reviewer is not only to criticize somebody else’s work but rather to help them to improve it with your feedback.</p><p>This last phrase is what I see as a major source of bad reviews.</p><blockquote>A bad reviewer often sees himself not as a peer of the authors with whom he or she wants to advance the state of research in his or her area, but as an ultimate (and sometimes superior) referee who is there to judge other’s work.</blockquote><p>The first approach to reviews <strong>takes time, requires patience </strong>and more than a pinch of <strong>goodwill</strong>. The second requires none of it and leads to a destructive half-random reviewing process where it can take years for important contributions to be actually published. Luckily, however, it is up to all of us to choose how we want it to be in the end.</p><h3>Afternote</h3><p>This article explains my approach to reviewing papers, but I am not the highest authority in this matter and I do not claim that it is the only right way to do it. There can be other opinions on how a good review should look like, as well as people who will find my reviews bad and uninformative. Also, there are different types of papers and reviewing a theoretical research paper may be very different from reviewing an applied research paper. The goal of this article was to show one possible way of how to do it hoping that it can be helpful for those who will find it suitable personally to them.</p><p>P.S. Thanks to Quentin Bouniot and Sofiane Dhouib for proofreading this article.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=f73bc037babc" width="1" height="1" alt=""><hr><p><a href="https://medium.com/data-science/reviewing-for-machine-learning-conferences-explained-f73bc037babc">Reviewing for Machine Learning Conferences Explained</a> was originally published in <a href="https://medium.com/data-science">TDS Archive</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Why worldwide collaboration can be up to  5 times more efficient than selfish behavior]]></title>
            <link>https://medium.com/data-science/why-worldwide-collaboration-can-be-up-to-5-times-more-efficient-than-selfish-behavior-b0a8e05ce50d?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/b0a8e05ce50d</guid>
            <category><![CDATA[game-theory]]></category>
            <category><![CDATA[covid19]]></category>
            <category><![CDATA[climate-change]]></category>
            <category><![CDATA[science]]></category>
            <category><![CDATA[economics]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Thu, 05 Nov 2020 14:09:01 GMT</pubDate>
            <atom:updated>2020-11-09T11:32:51.644Z</atom:updated>
            <content:encoded><![CDATA[<h3>Why worldwide collaboration can be up to 5 times more efficient than selfish behavior</h3><h4>Using game theory to understand the price of selfishness in times of pandemic and climate change</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*iR8eHnMK-isM6Tv_" /><figcaption>Photo by <a href="https://unsplash.com/@kylejglenn?utm_source=medium&amp;utm_medium=referral">Kyle Glenn</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><p>Have you ever wondered why we need all these international organizations that try to make all 195 countries of the world work together towards a common goal? You’ve probably heard from other people (or even presidents of certain countries) that it is just a waste of time and money, right?! Wouldn’t it be much easier if we just let everybody clean their mess?! Surprisingly, the consequences of opting for such an approach can be catastrophic on a global scale, and game theory provides us with a clear explanation of why this is the case.</p><h4><strong>Social optimum and selfish behavior</strong></h4><p>To understand the consequences of selfish behavior, I will use the example of the ongoing COVID-19 pandemic where the different actors (also called agents) are represented by independent countries, and the cost they are paying when facing the pandemic is quantified by the expenses (humanitarian or economic) required to contain it. To analyze the effect of collaboration in this gloom context, let‘s consider the following scenario:</p><ol><li>We assume that there is one country (for instance, that from which the virus originates) that pays the heaviest tribute. Let’s denote its overall loss by 1 (it can be 100 or 100K, 1 is used for the sake of simplicity).</li><li>Each country to which the virus spreads afterward can deal with it more efficiently when compared to the previous one. Let’s say that the second country can contain the pandemic by paying only half of what the first one has paid (1/2), while the third pays only one third (1/3), and so on. The final country number k will pay only 1/k fraction of the most affected country.</li><li>We assume that the <strong>socially optimal outcome</strong>, however, would have been to contain the pandemic in the first country by asking every existing country to <strong>participate equally</strong> in the expenses needed for that. Let’s say that this optimal outcome has a total cost of (1+a) where a &gt; 0 is a tiny overhead that accounts for the cost of mobilizing the required resources for the first suffering country.</li></ol><p>Note that while the first two assumptions do not reflect how everything happened in reality, the current situation can still be reduced to this model by sorting the countries based on their losses due to the pandemic in descending order.</p><p>It is straightforward to see that in the case of our model, the overall cost of all countries deciding not to collaborate is equal to the harmonic number</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/615/1*DYS4VNkzuwip9l69IGO1aw.png" /></figure><p>that behaves as shown in the plot below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*C43lLAZvkp-XmxgJ_7CbcA.png" /></figure><p>The difference between the selfish cost and the social optimum (denoted as collaborative baseline) is pretty huge, right?! You may wonder why anybody would opt for such an additional loss when having a much more efficient option at hand. Well, that is where game theory comes into play with the concept of Nash equilibrium, and to explain it, let us consider the different choices available to the agents in our game.</p><h4>Unilateral deviation</h4><p>Let us start with the last country from the example given above. In our model, this country is offered a choice to pay 1/k to contain the pandemic on its own or to pay (1+a)/k alongside other potential collaborators to achieve the overall socially optimal cost of 1+a.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/790/1*yPGxxLK0esp-cLxtfcAuWw.png" /></figure><p>When being selfish, this country will choose the option of paying 1/k to its benefit. Once this has happened, the country (k-1) faces a choice of paying (1+a)/(k-1) with others or opting for 1/(k-1) when dealing with the situation on its own.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/853/1*PUiL7PSvhrYjrYg3fy98UQ.png" /></figure><p>Once again, it chooses selfish behavior for the same reason as the country before. In the end, the selfish approach leads to the notion of <strong>Nash equilibrium</strong>: every country adapts the strategy that minimizes its costs, and no country can do better by unilaterally changing their strategy. The overall cost at the Nash equilibrium is then the harmonic number defined above.</p><h4><strong>Price of stability</strong></h4><p>The ratio between the cost at Nash equilibrium and the social optimum is commonly referred to as the<strong> Price of Stability </strong>(<strong>PoS</strong>).</p><blockquote><strong>Price of Stability</strong> quantifies the potential loss of efficiency between the best outcome of the selfish behavior (best Nash equilibrium) and the social optimum of a given game involving a set of strategic agents.</blockquote><p>In our case (the case of the so-called <strong>fair cost-sharing games</strong> studied by <em>Anshelevich et al., FOCS’04</em>), the <strong>PoS</strong> is bounded by the harmonic number defined above with k equal to 195 existing independent states in the world. If our very rough model of the pandemic is close to the truth, then the lack of worldwide collaboration can come at an alarming cost of</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/220/1*4GwhGo_UhObeXGGDMiQLSg.png" /></figure><p>Once again, if our model, despite its over-simplistic setting, is somewhere near to be correct then the lack of collaboration can cost <strong>up to 5 times more </strong>than the social optimum. Also note that this goes beyond my example with the COVID pandemic and extends to any situation where independent states have a choice between solving the problem on their own or opting for a world-wide collaboration. For instance, it highlights the importance of the Paris agreement signed by 195 countries to engage in the fight against climate change together and means, that we may well be avoiding the scenario of a drastic loss of efficiency with a possible devastating effect on our future.</p><h4>Fair cost sharing games of our life</h4><p>While my article may look like a simple illustration of an abstract game-theoretical result when applied to a situation of high-societal importance, its message is more general and destined for every one of us and our everyday choices.</p><p>Indeed, despite some seeming differences, all people on planet Earth play a game similar to that described above, albeit sometimes without even knowing it. Should I take the bus and share the travel cost with other passengers, or should I stick to using my car? Should I opt for the sharing economy platforms, or should I personally possess all the goods that I need? All these are fair cost-sharing games of our everyday life, and all of them are subject to the inefficiency that becomes more pronounced when more people are concerned with it.</p><p>To push this idea even further, it won’t be an exaggeration to say that we are often tempted to avoid individual sacrifices (the 1/k or (1+a)/k situation) for an immediate gain and thinking that one person’s action will not change much. But as my example shows, one person’s action can be enough to launch an unprecedented change to the behavior of the others. Indeed, if only the very first country were to choose to sacrifice its infinitesimal gain, it could have launched an incentive for others and make the collaborative alternative more attractive to them.</p><p>I agree that worldwide collaboration may not be easy at all. But the first step is to become conscious about its benefits. The next step is to think about it when you face any such choice in your life. And who knows maybe the step after that would be to enjoy a better world to live in.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=b0a8e05ce50d" width="1" height="1" alt=""><hr><p><a href="https://medium.com/data-science/why-worldwide-collaboration-can-be-up-to-5-times-more-efficient-than-selfish-behavior-b0a8e05ce50d">Why worldwide collaboration can be up to  5 times more efficient than selfish behavior</a> was originally published in <a href="https://medium.com/data-science">TDS Archive</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Hands-on guide to Python Optimal Transport toolbox: Part 2]]></title>
            <link>https://medium.com/data-science/hands-on-guide-to-python-optimal-transport-toolbox-part-2-783029a1f062?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/783029a1f062</guid>
            <category><![CDATA[machine-learning]]></category>
            <category><![CDATA[towards-data-science]]></category>
            <category><![CDATA[mathematics]]></category>
            <category><![CDATA[artificial-intelligence]]></category>
            <category><![CDATA[code]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Wed, 26 Aug 2020 13:40:27 GMT</pubDate>
            <atom:updated>2020-08-26T14:18:21.897Z</atom:updated>
            <content:encoded><![CDATA[<h4>Color transfer, Image editing and Automatic translation</h4><p>As a follow-up of my <a href="https://towardsdatascience.com/optimal-transport-a-hidden-gem-that-empowers-todays-machine-learning-2609bbf67e59">previous introductory article on optimal transport</a> and a first part of this guide provided by <a href="https://medium.com/@aboisbunon?source=post_page-----922a2e82e621----------------------">Aurelie Boisbunon</a> <a href="https://medium.com/@aboisbunon/hands-on-guide-to-python-optimal-transport-toolbox-part-1-922a2e82e621">here</a>, I will present below how you can solve different tasks with Optimal Transport (OT) in practice using the <a href="https://pythonot.github.io/">Python Optimal Transport (POT)</a> toolbox.</p><p>To start with, let us install POT using pip from the terminal by simply running</p><pre>pip3 install ot</pre><p>And voilà! If everything went well, you now have POT installed and ready to use on your computer. Let me now explain how you can reproduce the results from my previous article.</p><h4>Color transfer</h4><p>In this application our goal is to transfer the color style of one image onto another image in the smoothest way possible. To do this, we will follow <a href="https://pythonot.github.io/auto_examples/domain-adaptation/plot_otda_color_images.html#sphx-glr-auto-examples-domain-adaptation-plot-otda-color-images-py">the example</a> from the official webpage of the POT library and start by defining several supplementary functions needed when working with images:</p><pre>import numpy as np<br>import matplotlib.pylab as pl<br>import ot<br><br><br><a href="https://numpy.org/doc/stable/reference/random/legacy.html#numpy.random.RandomState">r</a> = <a href="https://numpy.org/doc/stable/reference/random/legacy.html#numpy.random.RandomState">np.random.RandomState</a>(42)<br><br><br>def im2mat(img):<br>    &quot;&quot;&quot;Converts an image to a matrix (one pixel per line)&quot;&quot;&quot;<br>    return img.reshape((img.shape[0] * img.shape[1], img.shape[2]))<br><br><br>def mat2im(X, shape):<br>    &quot;&quot;&quot;Converts a matrix back to an image&quot;&quot;&quot;<br>    return X.reshape(shape)</pre><p>So, first three lines here are just imports for numpy, matplotlib.pylab and ot packages. Then, we have two functions that allow us to convert an image represented by a 3d matrix (some people call them tensors) where the first dimension is the height of the image, the second one is its width, while the third is given by RGB coordinates of the pixels. Let’s now load some images to see what it means.</p><pre><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray">I1</a> = pl.imread(&#39;../../data/ocean_day.jpg&#39;).astype(<a href="https://docs.python.org/3/library/functions.html#float">np.float64</a>)/256<br><a href="https://numpy.org/doc/stable/reference/generated/numpy.ndarray.html#numpy.ndarray">I2</a> = pl.imread(&#39;../../data/ocean_sunset.jpg&#39;).astype(<a href="https://docs.python.org/3/library/functions.html#float">np.float64</a>)/256</pre><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F243JPx9kylZnAWCEWRmIw.png" /></figure><p>Here they are, the daytime ocean image and the sunset one provided directly in the POT toolbox. Note that originally all the pixel RGB coordinates are integers so that astype(<a href="https://docs.python.org/3/library/functions.html#float">np.float64</a>) converts them to floats. Then, each value is divided by 256 (the maximum value of each pixel’s coordinate) to normalize the data to lie in [0,1] interval. If we check their dimensions, we get the following</p><pre>print(I1[0,0,:])<br>[0.0234375 0.2421875 0.53125]</pre><p>This means that the first pixel in the bottom left corner has RGB coordinates given by the vector [R = 0.0234375, G= 0.2421875, B = 0.53125] (blue color dominates as expected for the daytime image). We now convert our tensors to a 2d matrix where each line is a pixel described by its RGB coordinates as follows:</p><pre>day = im2mat(I1)<br>sunset = im2mat(I2)</pre><p>Note that these matrices are rather large as can be seen by running the following code:</p><pre>print(day.shape)<br>(669000, 3)</pre><p>Let’s sample 1000 pixels randomly from each image to reduce the size of the matrices that we will apply OT to. We can do it as follows:</p><pre>nb = 1000<br>idx1 = r.randint(day.shape[0], size=(nb,)) <br>idx2 = r.randint(sunset.shape[0], size=(nb,))<br><br>Xs = day[idx1, :]<br>Xt = sunset[idx2, :]</pre><p>We now have two matrices with only 1000 rows and 3 columns in each of them. Let’s plot them in the RB (red-blue) plane to see the pixels of what color we actually sampled:</p><pre>plt.subplot(1, 2, 1)<br>plt.scatter(Xs[:, 0], Xs[:, 2], c=Xs)<br><em>#plt.axis([0, 1, 0, 1])<br></em>plt.xlabel(<strong>&#39;Red&#39;</strong>)<br>plt.ylabel(<strong>&#39;Blue&#39;</strong>)<br>plt.xticks([])<br>plt.yticks([])<br>plt.title(<strong>&#39;Day&#39;</strong>)<br><br>plt.subplot(1, 2, 2)<br><br>plt.scatter(Xt[:, 0], Xt[:, 2], c=Xt)<br><em>#plt.axis([0, 1, 0, 1])<br></em>plt.xlabel(<strong>&#39;Red&#39;</strong>)<br>plt.ylabel(<strong>&#39;Blue&#39;</strong>)<br>plt.title(<strong>&#39;Sunset&#39;</strong>)<br>plt.xticks([])<br>plt.yticks([])<br>plt.tight_layout()<em><br></em>plt.show()</pre><p>The result will look like this:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*mG3Y3rEKoyn-wMUiRy8kyA.png" /></figure><p>Everything seems to be set up to finally run our OT algorithm on them. To this end, we create an instance of the Monge-Kantorovich problem class and fit it on our images:</p><pre>ot_emd = ot.da.EMDTransport()<br>ot_emd.fit(Xs=Xs, Xt=Xt)</pre><p>Note that we create an instance of <em>ot.da.EMDTransport()</em> class that provides features for doing domain adaptation with OT and defines automatically the uniform empirical distributions (each pixel is a point having a probability 1/1000) and the cost matrix (squared Euclidean distance between vectors of pixel coordinates) when we call its <em>fit() </em>method. We can now “transport” one image onto another one using the coupling matrix as follows:</p><pre>transp_Xt_emd = ot_emd.inverse_transform(Xt=sunset)</pre><p>The function <em>inverse_transform()</em> we’ve just called transports the sunset image to the daytime one using a barycentric mapping: each transported pixel in the final result is an average of pixels from the sunset image weighted by the corresponding values of the coupling matrix. You can do the same thing the other way around by calling <em>transform(Xs=day) </em>too. We now plot the final result as follows:</p><pre>I2t = mat2im(transp_Xt_emd, I2.shape)<br><br>plt.figure()<br>plt.imshow(I2t)<br>plt.axis(<strong>&#39;off&#39;</strong>)<br>plt.title(<strong>&#39;Color transfer&#39;</strong>)<br>plt.tight_layout()<br>plt.show()</pre><p>And it gives the desired result:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*tlw_W6z3rVuz2CDVceGfKA.png" /></figure><h4><strong><em>Image editing</em></strong></h4><p>We now want to do a seamless copy that consists in editing an image by replacing a part of it using a patch of another image. For instance, this can be you face transported onto the fact of the Mona Lisa painting. To proceed, we will first need to download the <em>poissonblending.py</em> file from <a href="https://github.com/ncourty/PoissonGradient">this github repository</a>. Then, we will load three images from the data folder (you need to put them there beforehand) as follows:</p><pre><strong>import </strong>matplotlib.pyplot <strong>as </strong>plt<br><strong>from </strong>poissonblending <strong>import </strong>blend<br><br><br>img_mask = plt.imread(<strong>&#39;./data/me_mask_copy.png&#39;</strong>)<br>img_mask = img_mask[:,:,:3] <em># remove alpha<br><br></em>img_source = plt.imread(<strong>&#39;./data/me_blend_copy.jpg&#39;</strong>)<br>img_source = img_source[:,:,:3] <em># remove alpha<br><br></em>img_target = plt.imread(<strong>&#39;./data/target.png&#39;</strong>)<br>img_target = img_target[:,:,:3] <em># remove alpha</em></pre><p>First image is my portrait, second image provides the area of my portrait that will be copied into the Mona Lisa’s face. The pre-processing also removes the transparency and keeps only RGB values for each pixel. Overall, they will look as follows:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ANU13t3bvXHD8n0EevNqqQ.png" /><figcaption>You can adjust the mask using any image editor with simple geometrical objects.</figcaption></figure><p>The final result can then be obtained using the <em>blend()</em> function called as follows:</p><pre>nbsample = 500<br>off = (35,-15)</pre><pre>seamless_copy = blend(img_target, img_source, img_mask, reg=5, eta=1, nbsubsample=nbsample, offset=off, adapt=<strong>&#39;kernel&#39;</strong>)</pre><p>Once again, we apply OT only to a subset of 500 pixels as doing it for the whole image will take some time. The code behind this function involves many image pre-processing routines but what is of a special interest for us is the OT part. This is represented by the <em>adapt_Gradients_kernel() from</em> <em>poissonblending.py </em>that contains the following code:</p><pre>Xs, Xt = subsample(G_src,G_tgt,nb)<br><br>ot_mapping=ot.da.MappingTransport(mu=mu,eta=eta,bias=bias, max_inner_iter = 10,verbose=<strong>True</strong>, inner_tol=1e-06)<br>ot_mapping.fit(Xs=Xs,Xt=Xt)</pre><pre><strong>return </strong>ot_mapping.transform(Xs=G_src)</pre><p>The first line here extracts two samples of 500 pixels from the gradients G_src,G_tgt. Then, <em>ot.da.MappingTransport() </em>function learns a non-linear (kernelized) transformation that approximates the barycentric mapping that we have used in the previous example. You may wonder why is that needed? Well, the barycentric mapping relies on the coupling matrix that aligns only the samples it was fitted on (it’s shape is number of samples from the first distribution * number of samples from the second one) and thus it cannot be used to out-of-sample points. Finally, the return uses this approximation just as before to transport the gradient of my face onto the gradient of the Mona Lisa’s portrait. The final result is then given by:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DsDDxjH-J4v_Q8tTM1BXtw.png" /><figcaption>Rightmost image by Author.</figcaption></figure><h4>Automatic translation</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*FkzlFyi6qEVEUSeRXs5cIw.png" /></figure><p>For this last application, our goal is to take find an optimal alignment between the words in two sentences given in the different languages. As an example, we will work with the English proposition “<em>the cat sits on the mat</em>’’ and its French translation “<em>le chat est assis sur le tapis</em>’’ with the goal of recovering a matching that provides the correspondences “<em>cat</em>”- “<em>chat</em>”, “<em>sits</em>”- “<em>assis</em>” and “<em>mat</em>”- “<em>tapis</em>”. For this, we will need the nltk library that can be installed via pip as follows:</p><pre>pip3 install nltk</pre><p>We will also need to clone <a href="https://github.com/balikasg/WassersteinRetrieval">the following github repository</a> and to follow its readme in order to download the embeddings that will be used to describe our propositions (I also provide a shortcut to this here where you can find directly the embeddings for the considered pair).</p><p>Let us now do some usual imports and add two functions that will be used afterwards.</p><pre><strong>import </strong>numpy <strong>as </strong>np, sys, codecs<br><strong>import </strong>ot<br><strong>import </strong>nltk<br>nltk.download(<strong>&#39;stopwords&#39;</strong>) # download stopwords<br>nltk.download(<strong>&#39;punkt&#39;</strong>) # download punctuation <br><br><strong>from </strong>nltk <strong>import </strong>word_tokenize<br><strong>import </strong>matplotlib.pyplot <strong>as </strong>plt<br><br><strong>def </strong>load_embeddings(path, dimension):<br>    <em>&quot;&quot;&quot;<br>    Loads the embeddings from a file with word2vec format.<br>    The word2vec format is one line per words and its associated embedding.<br>    &quot;&quot;&quot;<br>    </em>f = codecs.open(path, encoding=<strong>&quot;utf8&quot;</strong>).read().splitlines()<br>    vectors = {}<br>    <strong>for </strong>i <strong>in </strong>f:<br>        elems = i.split()<br>        vectors[<strong>&quot; &quot;</strong>.join(elems[:-dimension])] =  <strong>&quot; &quot;</strong>.join(elems[-dimension:])<br>    <strong>return </strong>vectors<br><br><strong>def </strong>clean(embeddings_dico, corpus, vectors, language, stops, instances = 10000):<br>    <em><br>    </em>clean_corpus, clean_vectors, keys = [], {}, []<br>    words_we_want = set(embeddings_dico).difference(stops)<br>    <strong>for </strong>key, doc <strong>in </strong>enumerate(corpus):<br>        clean_doc = []<br>        words = word_tokenize(doc<em>)<br>        </em><strong>for </strong>word <strong>in </strong>words:<br>            word = word.lower()<br>            <strong>if </strong>word <strong>in </strong>words_we_want:<br>                clean_doc.append(word+<strong>&quot;__%s&quot;</strong>%language)<br>                clean_vectors[word+<strong>&quot;__%s&quot;</strong>%language] = np.array(vectors[word].split()).astype(np.float)<br>        <strong><br>        if </strong>len(clean_doc) &gt; 5 :<br>            keys.append(key)<br>        clean_corpus.append(<strong>&quot; &quot;</strong>.join(clean_doc))<br>    <strong>return </strong>clean_vectors</pre><p>First function is used to load the embeddings, while the second pre-process the text to remove all the stopwords and punctuation.</p><p>To proceed, we now load the embeddings for English and French languages as follows:</p><pre>vectors_en = load_embeddings(<strong>&quot;concept_net_1706.300.en&quot;</strong>, 300) <em><br></em>vectors_fr = load_embeddings(<strong>&quot;concept_net_1706.300.fr&quot;</strong>, 300) </pre><p>And define the two propositions to be translated:</p><pre>en = [<strong>&quot;the cat sits on the mat&quot;</strong>]<br>fr = [<strong>&quot;le chat est assis sur le tapis&quot;</strong>]</pre><p>Let us now clean our sentences as follows:</p><pre>clean_en = clean(set(vectors_en.keys()), en, vectors_en, <strong>&quot;en&quot;</strong>, set(nltk.corpus.stopwords.words(<strong>&quot;english&quot;</strong>)))</pre><pre>clean_fr = clean(set(vectors_fr.keys()), fr, vectors_fr, <strong>&quot;fr&quot;</strong>, set(nltk.corpus.stopwords.words(<strong>&quot;french&quot;</strong>)))</pre><p>This returns only the embeddings of meaningful words “<em>cat</em>”, “<em>sits</em>”, “<em>mat</em>” and “<em>chat</em>”, “<em>assis</em>” and “<em>tapis</em>”. Everything is now set for optimal transport to be used. As shown in the image above, we define two empirical uniform distributions over the terms and run OT between them with a cost matrix given by pairwise square Euclidean distances.</p><pre>emp_en = np.ones((len(en_emd),))/len(en_emd)<br>emp_fr =  np.ones((len(fr_emd),))/len(fr_emd)<br>M = ot.dist(en_emd,fr_emd)<br><br>coupling = ot.emd(emp_en, emp_fr, M)</pre><p>Once the coupling is obtained, we can now find a projection of our embeddings to 2d space with t-SNE and then plot the corresponding words and their matched pairs as follows:</p><pre>np.random.seed(2) # fix the seed for visualization purpose<br><br>en_embedded = TSNE(n_components=3).fit_transform(en_emd)<br>fr_embedded = TSNE(n_components=3).fit_transform(fr_emd)<br><br>f, ax = plt.subplots()<br>plt.tick_params(top=<strong>False</strong>, bottom=<strong>False</strong>, left=<strong>False</strong>, right=<strong>False</strong>, labelleft=<strong>False</strong>, labelbottom=<strong>False</strong>)<br>plt.axis(<strong>&#39;off&#39;</strong>)<br>ax.scatter(en_embedded[:,0], en_embedded[:,1], c = <strong>&#39;blue&#39;</strong>, s = 50, marker = <strong>&#39;s&#39;</strong>)<br><br><strong>for </strong>i <strong>in </strong>range(len(en)):<br>    ax.annotate(en[i], (en_embedded[i,0], en_embedded[i,1]), ha = <strong>&#39;right&#39;</strong>, va = <strong>&#39;bottom&#39;</strong>, fontsize = 30)<br><br>ax.scatter(fr_embedded[:,0], fr_embedded[:,1], c = <strong>&#39;red&#39;</strong>, s = 50, marker = <strong>&#39;s&#39;</strong>)<br><br><strong>for </strong>i <strong>in </strong>range(len(fr)):<br>    ax.annotate(fr[i], (fr_embedded[i,0], fr_embedded[i,1]), va = <strong>&#39;top&#39;</strong>, fontsize = 30)<br><br>coupling /= np.max(coupling)<br><br><strong>for </strong>i, j <strong>in </strong>enumerate(np.array(np.argmax(coupling, axis= 1)).flatten()):<br>    ax.plot([en_embedded[i, 0], fr_embedded[j, 0]], [en_embedded[i, 1], fr_embedded[j, 1]], c=<strong>&#39;k&#39;</strong>)<br><br>plt.show()</pre><p>This gives the final result:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*-Ih9cJ-zWRaB9QWvtNo2-Q.png" /><figcaption>Image by Author.</figcaption></figure><p>Note that the latter examples can be extended to do automatic translation based on the wikipidea data available on the github repository from which we took the initial code. For more details, you can also check the corresponding paper and the results therein to further boost the performance.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=783029a1f062" width="1" height="1" alt=""><hr><p><a href="https://medium.com/data-science/hands-on-guide-to-python-optimal-transport-toolbox-part-2-783029a1f062">Hands-on guide to Python Optimal Transport toolbox: Part 2</a> was originally published in <a href="https://medium.com/data-science">TDS Archive</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[What I learned from hitchhiking]]></title>
            <link>https://medium.com/illumination/what-i-learned-from-hitchhiking-9627c9676f7d?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/9627c9676f7d</guid>
            <category><![CDATA[hitchhiking]]></category>
            <category><![CDATA[travel]]></category>
            <category><![CDATA[lessons-learned]]></category>
            <category><![CDATA[experience]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Mon, 15 Jun 2020 14:33:19 GMT</pubDate>
            <atom:updated>2020-06-15T14:33:19.763Z</atom:updated>
            <content:encoded><![CDATA[<h4>30,000kms on the thumb around the globe in 1000 words</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*1nxFxN_EvIuP6frL" /><figcaption>Photo by <a href="https://unsplash.com/@bbergher?utm_source=medium&amp;utm_medium=referral">Bruno Bergher</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><p>In the last 8 years of my life, I hitchhiked in almost 30 different countries with a total of around 30,000kms traveled. While I do not hitchhike that much nowadays, I feel that the experiences gained on the road learned me a lot about life and people’s behavior in general, and sharing the lessons I have learned from them with you is what I would like to do below.</p><h4>Lesson #1: Create your chances</h4><blockquote>Getting lucky is a matter of you doing something for it to happen.</blockquote><p>I learned this one from all the times when I got dropped off in places from which it seemed impossible to get a ride. I was standing on the side of the road watching all the cars passing by so fast they couldn’t probably even see me, and understanding that I needed to get extremely lucky to get out of there. You know what my solution was in these situations? I usually chose whatever direction felt safe, put my backpack on the shoulders, and started advancing in that direction to improve my current position rather than staying immobile and waiting for something to happen. I do not know why and how this works exactly, but cars that some minutes ago were speeding away now were making their best to pull over safely to pick me up. When asked why they stopped, the drivers usually replied “Oh, I just saw you walking.”<br>My hint about this one is that people tend to have more sympathy when they see that you are trying to find a solution, and not just hoping for it to come as a blessing or a stroke of good luck. Every time you are cornered by the unfavorable circumstances, try to seek this little tiny improvement, this extra 1% increasing your chances to get out of the situation you are in. Somehow, this one extra percent is often exactly what you need to succeed, whether it will be on your own or with the help of somebody else.</p><h4><strong>Lesson #2: Know what you want</strong></h4><blockquote>You need to decline offers that will not help you to advance in the long-term.</blockquote><p>You may be hitchhiking from a great spot heading somewhere far, and suddenly it starts to rain heavily. You are getting all wet and angry, and then a car stops by and offers to get you to a nearby town. You may be tempted to go with it as being inside a car is great when it rains outside, but you know that this town is not exactly on the way to where you are heading. Should you accept or decline?</p><p>From my experience, settling on a more attractive short-term goal will eventually lead to a waste of time in the long-term if the two are not aligned. For instance, consider the following situation. You quit your job as you feel ready to move on and aim for a bigger project. Then, you receive an offer for a slightly better job, with a bigger salary, but still not matching your current ambitions. Accepting it may be tempting in the short-term due to obvious advantages it brings immediately, but if you do so, you may waste your time and energy without advancing towards what was important to you in the first place. Keeping the final goal in mind is something that helps to remain focused and have a broader perspective when making important decisions.</p><h4>Lesson #3: Do not expect too much</h4><blockquote>Tame your expectations whether they are too optimistic or too pessimistic.</blockquote><p>Hitchhiking is one of those things that attract people with its adventurous allure. When you first put your feet on the side of the road, you think about all the crazy adventures that happened to the heroes of your favorite Jack Kerouac’s novel and hope that the same will happen to you too. The reality, however, is somewhat different and is full of long waits and uneventful encounters, with only one out of a hundred rides ending up in something memorable. In everyday life, the situation is quite similar when we tend to expect too much from any new person we meet, any trip we embark on, and, in general, from anything that happens to us.</p><p>There is a very intuitive mathematical explanation for high expectations leading to unhappiness that I give to my students, and it is the Bayes’ theorem. It roughly tells you that the probability of having a certain outcome given fixed circumstances depends on your prior beliefs. If your beliefs are biased (either too pessimistic or too optimistic), then you will often be disappointed as your expectations will rarely match the reality. Hitchhiking taught me that unique experiences just happen out of the blue and that enjoying the ride in-between them, however good, bad or boring it can be, is the best way of having the most of it all.</p><h4>Lesson #4: Listen to people</h4><blockquote>Learn to understand what is your role and what is expected from you.</blockquote><p>I was a very shy guy back in my student days, and I didn’t talk too much with people that I used to meet. For obvious reasons, hitchhiking changed me quite a lot in this sense, but what I value most about it is not my current capacity of talking to strangers, but the ability to understand what people want from me. Let me explain it with the following example. You get into the car, exchange a couple of introductory phrases with the driver and then the conversation comes to a stall. What should you do now? When I was starting to hitchhike, I thought that my role was to entertain the driver and any such moments of uncomfortable silence made me think that I was not fulfilling my mission. It took me some time to understand that not all people expect the same from me and failing to recognize their expectations and force you way around them can make you miss a lot.</p><p>One of the most memorable rides that I got was from a Canadian guy who drove me from Halifax airport to North Sydney’s ferry terminal. From the moment I got in the car, I understood that he needed somebody to talk to and was not interested in who I was or what was there to see in my home-country. So I barely opened my mouth during the ride and I listed about his life as if reading a book: he was a very good story-teller, and his life was worth hearing about. The same can go the other way around as well. Some people only wait for you to take the lead of the conversation, while others are absorbed in their thoughts and do not want to be engaged in any sort of exchange at all.</p><p>This is a hard lesson to learn as all people are different and we are no telepaths and cannot read their minds. The only thing that we can do, however, is being patient and willing to understand how to create the best synergy with the person in front of us. It may seem little, but sometimes it can go a long way.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=9627c9676f7d" width="1" height="1" alt=""><hr><p><a href="https://medium.com/illumination/what-i-learned-from-hitchhiking-9627c9676f7d">What I learned from hitchhiking</a> was originally published in <a href="https://medium.com/illumination">ILLUMINATION</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Optimal transport: a hidden gem that empowers today’s machine learning]]></title>
            <link>https://medium.com/data-science/optimal-transport-a-hidden-gem-that-empowers-todays-machine-learning-2609bbf67e59?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/2609bbf67e59</guid>
            <category><![CDATA[artificial-intelligence]]></category>
            <category><![CDATA[mathematics]]></category>
            <category><![CDATA[machine-learning]]></category>
            <category><![CDATA[towards-data-science]]></category>
            <category><![CDATA[computer-science]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Mon, 15 Jun 2020 11:05:55 GMT</pubDate>
            <atom:updated>2020-08-26T13:49:21.011Z</atom:updated>
            <content:encoded><![CDATA[<h4>Explaining one of the most emerging methods in machine learning right now</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*qDMcdbjudSqdSf-s0i7S7w.png" /><figcaption>Source: <a href="https://perso.liris.cnrs.fr/nicolas.bonneel/spot/">Nicolas Bonneel</a>, via <a href="https://www.youtube.com/watch?time_continue=2&amp;v=KTFn3YWN7b4&amp;feature=emb_logo">Youtube</a></figcaption></figure><p>Would you believe me if I were to say that there is a single solution to such different problems as brain decoding in neuroscience, shape reconstruction in computer graphics, color transfer in computer vision, and automatic translation in neural language processing? And if I were to add <a href="https://towardsdatascience.com/why-transfer-learning-works-or-fails-27dcb8095670">transfer learning</a>, image registration, and spectral unmixing for musical data to this list and the fact that the solution I am talking about has nothing to do with deep learning? Well, if you couldn’t guess the right answer, then keep on reading as this article is about Optimal Transport (OT): a mathematical theory dating back to the late 18th century that has flourished recently in both pure mathematics (2 Fields medals in last 12 years!) and machine learning to become one of the most emerging topics to learn about right now. Join me on this walk through the beautiful history of OT spanning the last three centuries and learn why it became such an important part of today’s machine learning.</p><h4>The beginning</h4><p>Like many good things, it all started in France in the late 18th century under the rule of Louis XVI, around a decade before the French revolution of 1789. On that sunny day, Louis XVI looked concerned when discussing with one of the most prominent scientists of his country, Gaspard Monge, a matter of the highest governmental importance.</p><p>“You see, Gaspard,” said Louis XVI, “we have three types of cheese certification in France: fermier cheese fully made without leaving the farm, artisanal cheese made from milk from the nearby farms by artisans, and laitier cheese made from milk from the same region at the factories. The problem is that artisans making artisanal cheese tend to waste too much time riding around and picking up milk from the farms that are too far away. I would like you, Gaspard, to help us to deal with the issue by figuring out which farm gives all of its milk to which artisan so that the cheese is abundant and the life is pretty.</p><p>“Oh!” exclaimed Gaspard somewhat excitedly, “it’s like this problem with ore and mines that I’ve been working on recently with my fellow metallurgists.”</p><p>“Whatever, Gaspard,” replied his Majesty. “As long as the cheese is on the table and the people are happy.”</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/923/1*ek4Fj8tlYMQmNF0mel4kEQ.png" /><figcaption>Assuming <strong>one cow’s milk</strong> is enough to produce <strong>one cheese</strong>, the optimal <strong>T</strong> here would be an <strong>assignment</strong> given by the arrows with an overall distance of 19kms.</figcaption></figure><p>After this conversation with the king, Gaspard immediately figured out that the problem he is dealing with can be expressed using the table on the left. Here, each cell gives the distance between the farms and the artisans, while the number of cows and cheeses indicates their respective production capacities and supply needs. The goal then would be to find a function <strong>T, </strong>called the <strong>Monge map</strong>, that assigns each farm to an artisan in an optimal way by taking into account both the distances between them and their respective demands.</p><p>Despite its simplistic appearance, it took mathematicians almost two hundred years to fully characterize this problem (some 140 years less than the famous Fermat’s Last Theorem), until Yann Brenier showed in 1991 that it admits a solution in the general (continuous) case for some common distances used in science. But, before that, the following important thing happened.</p><h4>The unexpected benefits of communism</h4><p>Bookkeeper Nina, whose job in the late 1930s was what Google search is for us today, was staring at the shelves of the Leningrad State Library trying to find the Charles Baudelaire’s book <em>“Les Fleurs du mal”. </em>The latter was asked for by Leonid, a poetry-lover in his late twenties now waiting impatiently at the counter.</p><p>“We do not have Charles Baudelaire, comrade Kantorovich,” she said when coming back to her workplace. “But here is a nice article about metallurgy that you, as an inventor of linear programming and an honorable Soviet citizen, may find interesting”.</p><p>The nice article she referred to was Gaspard Monge’s “<em>Mémoire sur la théorie des déblais et des remblais” </em>that Leonid Kantorovich took somewhat disappointingly and, while reading it back at home, figured out how it can be improved.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/923/1*PM1jFdICdSMWYRD9O3HFfg.png" /><figcaption>In Kantorovich formulation, we are allowed to <strong>split the production</strong> among the artisans to minimize the distances covered by them.</figcaption></figure><p>“The issue with <strong>Monge problem</strong> is that those monarchists do not share stuff, while we, Soviet people, share everything with each other,” said Leonid to himself. “Take milk, for instance. Why would we care about not allowing farms to split its production among all factories?”</p><p>This brought him to the idea that would later be called the <strong>Monge-Kantorovich problem </strong>where contrary to seeking for a function <strong>T</strong> assigning each farm to one artisan, we now want to find a (probabilistic) function <strong>Г</strong> that is allowed to split the production of farms among all artisans.</p><p>This problem, contrary to the Monge’s one, can be shown to 1) always have a solution under some mild assumptions regarding the distance used and 2) earn Nobel Prize, Fields medals and other distinguished awards for a dozen of different researchers working on it in the last 60 years. Finally, it hints on why all the cheese in post-Soviet countries tastes the same despite being sold under different names and that, my friend, is a great modern mystery to put some light on (believe me, I was born in Ukraine).</p><h4><strong>Low-level view</strong></h4><p>As you may have guessed already, all the stories presented so far are simplified (and even made up) for entertainment sake and to provide a general intuition behind OT as the true problems considered by both Monge and Kantorovich were much more complicated than that.</p><blockquote>In fact, the<strong> true power of OT</strong> is hidden in its ability to act as a mechanism of <strong>transforming</strong> <strong>one</strong> (continuous) <strong>probability distribution</strong> <strong>into another one</strong> with the <strong>lowest possible effort.</strong></blockquote><p>The latter criterium is measured by some function providing a dissimilarity for any pair of points from the two distributions.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*MUaR6OETeWd4UymBk0O7fQ.png" /><figcaption>Monge problem with discrete probability distributions represented by histograms and their fitted continuous approximations given by green and purple lines.</figcaption></figure><p>Getting back to my example with farms and artisans, I can now define two discrete distributions over them with probabilities given by the production of each farm (3 cows out of 7 give me the probability of 3/7 ) and the demand of each artisan. I will then obtain two histograms, and the distances considered before would become the costs of transforming the points from my first distribution into those from the second one.</p><p>“But why is that so important?” you may ask.</p><blockquote>Well, mainly because a <strong>whole awful lot of objects</strong> manipulated by data science practitioners <strong>can be modeled as probability distributions</strong>.</blockquote><p>For instance, any image can be seen as a distribution over pixels with probabilities given by their intensities normalized to sum to one. The same holds for text documents that you can see as discrete distributions of words with probabilities given by their frequencies of occurrence. The examples are countless, and OT allows you to deal with all of them whenever you need to compare two such complex objects.</p><p>“But we usually use distances to compare objects!” you may rightfully notice.</p><blockquote>Well, OT got it all covered once again as <strong>Kantorovich formulation</strong> of the problem defines a <strong>true metric</strong> on the space of distributions, called the <strong>Wasserstein distance</strong>, that obeys the triangle inequality and vanishes when two distributions are equal.</blockquote><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*HPVUdXzgSMheg1YZUX_kCA.png" /><figcaption>Interpolation between Monge and Kantorovich portraits using OT. The intermediate images show the shortest path for transforming one image into another.</figcaption></figure><p>Moreover, as OT problem finds the most efficient way of transforming one distribution into another, you can use its solution to interpolate smoothly between them and obtain intermediate transformations along this path. To illustrate this, let’s take the image of Gaspard Monge and that of Leonid Kantorovich and solve the OT problem between them. The result in the figure above shows a geodesic path between the two images given by the most efficient way (that of the shortest path) of gradually transforming one image into another with the geometry (or pairwise distances between the pixels in our case) specified by the cost function you are using. Pretty cool, heh?! Let’s now finally get back to the modern times to see what exactly this means for machine learning.</p><h4>Applications’ trailer</h4><p>Below, I will present only some applications of OT in computer vision and neural language processing with the code used to produce the results covered in detail in <a href="https://towardsdatascience.com/hands-on-guide-to-python-optimal-transport-toolbox-part-2-783029a1f062">“Hands-on guide to Python Optimal Transport toolbox: Part 2”</a>. Also, check out the introduction to Python Optimal Transport toolbox in this article <a href="https://medium.com/@aboisbunon/hands-on-guide-to-python-optimal-transport-toolbox-part-1-922a2e82e621">“Hands-on guide to Python Optimal Transport toolbox: Part 1”</a>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*RDc5ZWTVJmUgOh5JW27qIA.png" /><figcaption><strong>Top row</strong>: <strong>left </strong>— original day sky image; <strong>right </strong>— original sunset sky image; <strong>middle</strong> — the result of color transfer. <strong>Bottom row</strong>: the coordinates of point in the Blue and red frequencies highlighting the different color styles of the two images. The reader may use <a href="https://pythonot.github.io/auto_examples/domain-adaptation/plot_otda_color_images.html#sphx-glr-auto-examples-domain-adaptation-plot-otda-color-images-py">the following code</a> to reproduce the results.</figcaption></figure><p><em>Color transfer. </em>In <a href="https://arxiv.org/pdf/1307.5551.pdf">this application</a>, we have two images: one with a blue sky over the ocean and one with the sunset. Our goal is to transfer the color style of the first image to the second one so that the colors of the sunset image will look identical to that of the daytime one. As before, we consider both images as probability distributions over pixels given by 3-dimensional vectors in the RGB cube. We sample 1000 pixels from each image and align them using OT. The final result is given in the middle image on the left and shows a highly realistic sunset image with the daytime colors.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*DsDDxjH-J4v_Q8tTM1BXtw.png" /><figcaption>A result of color gradient adaptation with OT for Poisson image editing using the code from <a href="https://github.com/ncourty/PoissonGradient">here</a>. Note that the final result can be improved by using a more precise image mask.</figcaption></figure><p><em>Image editing. </em>For this task, the goal is to edit an image by replacing a part of it using a patch of another image. For instance, in the image on the left, you can see my face, the Mona Lisa’s face, and the result of a seamless copy of my face on hers. Optimal transport here is applied to color gradients of the two images, and then the Poisson equation is solved to calculate the edited image. For more cool examples of this, check <a href="https://papers.nips.cc/paper/6312-mapping-estimation-for-discrete-optimal-transport.pdf">this paper</a>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*FkzlFyi6qEVEUSeRXs5cIw.png" /><figcaption>Aligning the embeddings of documents written in different languages using OT provides state-of-the-art results for <a href="https://arxiv.org/pdf/1809.00013.pdf">automatic translation</a> and <a href="https://arxiv.org/pdf/1805.04437.pdf">cross-lingual information retrieval</a>.</figcaption></figure><p><em>Automatic translation</em>. Let’s now turn our attention to something else than images and consider text corpora. Our goal would be to find a matching of words and phrases in two different languages. Let take the following pair as an example: given English proposition “<em>the cat sits on the mat</em>’’ and its French translation “<em>le chat est assis sur le tapis</em>’’, we would like to find a matching that provides the correspondences “<em>cat</em>”- “<em>chat</em>”, “<em>sits</em>”- “<em>assis</em>” and “<em>mat</em>”- “<em>tapis</em>”. You see where I am heading with this one, don’t you?! I will define a distribution over each proposition and treat each word as a single point. I will then use the distances between their embeddings to match the two distributions with OT. The result of this matching can be seen on the left.</p><p>There are many more applications where OT has shown to be useful, including, for instance, <a href="https://perso.liris.cnrs.fr/julie.digne/articles/jmiv.pdf">shape reconstruction,</a> <a href="https://arxiv.org/pdf/1506.05439.pdf">multi-label classification</a>, <a href="https://arxiv.org/pdf/1701.07875.pdf">GAN training</a>, <a href="https://arxiv.org/pdf/1503.08596.pdf">brain decoding</a>, <a href="http://openaccess.thecvf.com/content_cvpr_2015/papers/Kolouri_Transport-Based_Single_Frame_2015_CVPR_paper.pdf">image super-resolution</a>, and many, many others. OT is a huge universe in itself, and it only keeps on rising, as confirmed by an increasing number of papers studying it at top machine learning conferences.</p><h4><strong>Afternote</strong></h4><p>This article deliberately skips all the mathematical details on the Monge and Monge-Kantorovich problems, the duality theory, and the recently introduced entropic regularization to remain accessible even for the readers without mathematical background. For more technical details, the reader may refer to many beautiful overviews of the OT theory, including a <a href="https://arxiv.org/pdf/1803.00567.pdf">recent book</a> by Gabriel Peyré and Marco Cuturi and a much more in-depth <a href="https://ljk.imag.fr/membres/Emmanuel.Maitre/lib/exe/fetch.php?media=b07.stflour.pdf">monograph</a> by Cédric Villani.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=2609bbf67e59" width="1" height="1" alt=""><hr><p><a href="https://medium.com/data-science/optimal-transport-a-hidden-gem-that-empowers-todays-machine-learning-2609bbf67e59">Optimal transport: a hidden gem that empowers today’s machine learning</a> was originally published in <a href="https://medium.com/data-science">TDS Archive</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Why transfer learning works or fails?]]></title>
            <link>https://medium.com/data-science/why-transfer-learning-works-or-fails-27dcb8095670?source=rss-fde5e3dd8903------2</link>
            <guid isPermaLink="false">https://medium.com/p/27dcb8095670</guid>
            <category><![CDATA[domain-adaptation]]></category>
            <category><![CDATA[transfer-learning]]></category>
            <category><![CDATA[artificial-intelligence]]></category>
            <category><![CDATA[machine-learning]]></category>
            <category><![CDATA[towards-data-science]]></category>
            <dc:creator><![CDATA[Ievgen Redko]]></dc:creator>
            <pubDate>Wed, 13 May 2020 14:32:36 GMT</pubDate>
            <atom:updated>2020-07-03T15:34:14.654Z</atom:updated>
            <content:encoded><![CDATA[<h4>An (almost) math-free guide to understanding the theory behind transfer learning and domain adaptation.</h4><figure><img alt="" src="https://cdn-images-1.medium.com/max/888/1*jiJm0xdwOq3lewAAIw9amw.jpeg" /><figcaption>source: Sebastian Ruder, via <a href="https://www.slideshare.net/SebastianRuder/transfer-learning-the-next-frontier-for-machine-learning">slideshare</a></figcaption></figure><p>During the NIPS tutorial talk given in 2016, Andrew Ng said that <strong>transfer learning</strong> — a subarea of machine learning where the model is learned and then deployed in <em>related, yet different, areas</em> — will be the next driver of machine learning commercial success in the years to come. This statement would be hard to contest as avoiding learning large-scale models from scratch would significantly reduce the high computational and annotation efforts required for it and save data science practitioners lots of time, energy, and, ultimately, money.</p><p>As an illustration of these latter words, consider <a href="https://research.fb.com/wp-content/uploads/2016/11/deepface-closing-the-gap-to-human-level-performance-in-face-verification.pdf">Facebook’s DeepFace</a> algorithm that was the first to achieve a near-human performance in face verification back in 2014. The neural network behind it was trained on <strong>4.4 million</strong> <strong>labeled</strong> faces — an overwhelming amount of data that had to be collected, annotated, and then trained on for 3 full days without taking into account the time needed for fine-tuning. It won’t be an exaggeration to say that most of the companies and research teams without Facebook’s resources and deep learning engineers would have to put in months or even years of work to complete such a feat, with most of this time spent on collecting an annotated sample large enough to build such an accurate classifier.</p><p>This is where transfer learning magically steps in by allowing us to use the same model across related datasets just as we would have done it if they were to come from the same source. Despite being quite efficient and helpful for such challenging tasks as computer vision and natural language processing, transfer learning algorithms also fail badly in practice, and explaining why it may or may not happen is what I will attempt to do below.</p><h4>Getting back to the roots</h4><p>To start my brief and painless introduction to transfer learning theory, let me introduce Homer, a guy in his late 30s who got all excited because of the hype around machine learning and decided to automatically classify all the weird stuff he buys on Aliexpress for his online shop. The main motivation of Homer stemmed from his laziness, the fact that translated English descriptions on Aliexpress were usually quite confusing, to say the least, meaning that only photos of what Homer buys were providing any information about the actual item.</p><p>And so, Homer downloaded a huge annotated dataset of items sold on Amazon, hoping that a classifier learned on them was going to work well on his images from Aliexpress too. “What made him think so?” you may ask. Well, first of all, Homer assumed that with that many images from Amazon, he can learn a low-error classifier for them using a state-of-the-art deep neural network with 1 trillion layers. Also, during all summer vacations spent at his grandmother’s house near the sea, Homer had a chance to read <a href="https://en.wikipedia.org/wiki/Vapnik%E2%80%93Chervonenkis_dimension#In_statistical_learning_theory">all the latest work</a> of Mr. Vapnik and Mr. Chervonenkis who, (very) broadly speaking, had suggested the following inequality:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*ANB9Xx8nglcToLrDPRHHXg.png" /></figure><p>Homer knew that the first term on the right-hand side can be made as small as desired due to neural networks’ capacity to <a href="https://arxiv.org/pdf/1611.03530.pdf">learn well from any rubbish</a> fed into them. Also, Homer supposed that the high complexity of the neural network in the numerator of the second term was going to be compensated by the large sample size at his disposal, thus bringing it close to 0 too. The last piece of the puzzle that bothered Homer is the left-hand side, as he was not sure whether the classification error on unseen items from Amazon was going to be close to that achieved on items from Aliexpress. To deal with this, he made the following simple assumption:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*9l_kwv5gGZJ6maOmX1o9FQ.png" /></figure><p>“What kind of distance?” a curious reader will ask and will be right to do so. But Homer didn’t not care about such details and, being now quite happy with himself, proceeded to the following ultimate inequality:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*zubqPPqIwWR9KpiPAgNtXg.png" /></figure><blockquote>“Now I know what to do,” said Homer to himself, meaning transfer learning and not his unsettled life in general. “First, I need to find a way to transform images from Amazon so that they will look as similar as possible to those from Aliexpress, thus reducing the distance between them. Then, I will learn a low-error classifier on the transformed images, as I still have ground-truth labels for them, and will apply this classifier further on my images from Aliexpress.”</blockquote><p>After some moments of thinking, he became uncertain about his idea. “Is there is something I am missing here?” he asked himself, while ordering a laser saber umbrella from Aliexpress’ website, and as it happens, he did.</p><h4>What is supposed to be called similar?</h4><p>While Homer’s intuition about transfer learning formalized in the last inequality was generally right, he still lacked a well-defined notion of a distance that he could have used as a measure of transferability between two datasets.</p><blockquote>“Roughly speaking,” Homer reasoned with himself, “ there are two possible ways of comparing datasets: an unsupervised and a supervised one. If I go with the supervised one, it means that I take into account both images and their labels to measure the distance; if I opt for the unsupervised one, I consider images only.”</blockquote><p>Both these approaches bothered Homer, but for different reasons. For the supervised approach, he had to have labels for Aliexpress images, which was something that he was trying to obtain with transfer learning in the first place. As for the unsupervised one, he thought that it was not accurate enough as two images can seem similar even when they belong to different categories. “How’s that can be?” you may wonder. Well, just take a look at the following two items sold on <a href="https://www.aliexpress.com/item/32582734205.html">Aliexpress</a> and <a href="https://www.amazon.co.uk/1-5kg-Fresh-Salmon-Fillet-Boned/dp/B0747NCFJB">Amazon</a> websites and tell me to which category each of them belongs.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*3475GU2nzcoNLPYx6Sf2cQ.jpeg" /><figcaption>An image of an item sold on Aliexpress (left, source: Ningbo Creight Co., via aliexpress) and on Amazon (right, Regal Fish Supplies, via amazon).</figcaption></figure><p>Obviously, the one on the left is a sleeping pillow (it was obvious, wasn’t it?!), while that on the right is a salmon fillet. To avoid this sort of confusion, Homer decided to put forward the following assumption:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*F-zLQ7BU2DZ7HA7HpoSvUQ.png" /></figure><p>Homer thought that putting the labeling functions into the equation— functions that output the category of any possible item from their respective online platform — is the most straightforward way to account for both the annotations of the datasets and the actual similarity of their images. Surprisingly, this is how he (almost) came up with a result extremely close to Theorem 1 from <a href="http://www.alexkulesza.com/pubs/adapt_mlj10.pdf">the seminal paper on transfer learning theory</a>.</p><h4>One theory to bind them all</h4><p>You may be quite surprised if I tell you that most of the papers on transfer learning theory boil down to the inequality derived by Homer from his very basic understanding of machine learning principles. The only difference between <a href="https://arxiv.org/pdf/2004.11829v1.pdf">those numerous papers</a> and Homer’s down-to-earth reasoning is that Homer had to get through by making the assumptions and not by actually proving the desired result contrary to the published works.</p><p>I will now show you a slightly reformulated inequality that captures the essence of both Theorem 1 and Homer’s reasoning, and then we will see how it can be used to justify all those <a href="https://towardsdatascience.com/search?q=transfer%20learning">transfer learning algorithms that were discussed on Medium before</a>. My reformulation writes as follows:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*WG2L3TY0I02CnFlwrid44Q.png" /></figure><p>Here, my target domain is any dataset that I may want to categorize without manually annotating it; in Homer’s example, it consists of images of the items sold on Aliexpress. My source domain, abundant Amazon images in that same example, is any annotated dataset for which I can produce a low-error model used in the target domain afterwords. The second term on the right-hand side is an unsupervised distance between the two domains that we can usually calculate without knowing the labels of instances in any of them. Researchers working on transfer learning proposed many different candidates for this term, and most of them took the form a certain divergence between the (marginal) distributions of the two domains. Finally, the third term represents what is usually called <em>the a priori adaptability</em>: a non-estimable quantity that we can compute only when the true target domain’s labeling function is known. This latter observation brings us to the following important conclusion.</p><blockquote>While a transfer learning algorithm can explicitly minimize the first two terms of our inequality, the a priori adaptability term remains <strong>uncontrollable</strong>, potentially leading to a <strong>failure of transfer learning</strong>.</blockquote><p>If you wait for a magic solution at this point, then I will have to disappoint you by saying that there is no such solution. You can use <a href="https://towardsdatascience.com/deep-domain-adaptation-in-computer-vision-8da398d3167f">kernel-based, moment matching, or adversarial approaches</a>, but it won’t change anything: in the end you will be left at the mercy of the non-estimable term that may ultimately impact the final performance of your model in the target domain with its invisible hand. The good news, however, is that in most cases, it will still work better than doing no transfer at all.</p><h4><strong>Back to real life</strong></h4><p>I will now present a simple example provided in <a href="http://proceedings.mlr.press/v9/david10a/david10a.pdf">this paper</a> that highlights one of the pitfalls of transfer learning approach following the philosophy described above. In this example, we will consider two 1-dimensional datasets, representing the source and target domains, and generated using the code given below.</p><pre><strong>import </strong>numpy <strong>as </strong>np<br><br>np.set_printoptions(precision=1)</pre><pre><strong>def </strong>generate_source_target(xi):<br></pre><pre>    k=0<br>    size = int(1./(2*xi))<br>    source = np.zeros((size,))<br>    target = np.zeros_like(source)<br><br>    <strong>while </strong>(2*k+1)*xi&lt;=1:<br>        source[k] = 2*k*xi<br>        target[k] = (2*k+1)*xi<br>        k+=1<br>    <strong>return </strong>source, target</pre><p>Executing this piece of code produces two sets of points in the interval [0,1], as in the figure below.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sOVmKezaTVbtqSZIaC_XFQ.png" /></figure><p>For instance, when ξ = 0.1, it will return the following two lists:</p><pre>source,target = generate_source_target(1./10)</pre><pre>print(source)<br>[0. 0.2 0.4 0.6 0.8]</pre><pre>print(target)<br>[0.1 0.3 0.5 0.7 0.9]</pre><p>I will further attribute label 1 to all points from the source domain and label 0 to those from the target domain. My final learning samples will thus become:</p><pre>source_sample = [(round(i,1),1) <strong>for </strong>i <strong>in </strong>source]<br>target_sample = [(round(i,1),0) <strong>for </strong>i <strong>in </strong>target]<br><br>print(source_sample)<br>[(0.0, 1), (0.2, 1), (0.4, 1), (0.6, 1), (0.8, 1)]</pre><pre>print(target_sample)<br>[(0.1, 0), (0.3, 0), (0.5, 0), (0.7, 0), (0.9, 0)]</pre><p>Is it easy to find a perfect classifier for each of these samples? Yes, it is, as, in case of the source domain, it can be done using a threshold function that outputs 1 for points whose coordinate is smaller than 0.8, and 0, otherwise. The same holds for the target domain, but this time the classifier will output 0 for all points whose coordinate is smaller than 0.9, and 1, otherwise. Finally, we will now make sure that the distance between these two domains depends on ξ and thus can be made arbitrarily small by reducing it. To do this, we will use the 1-Wasserstein distance that is exactly equal to ξ in our case. Let’s verify it using the following code:</p><pre><strong>from </strong>scipy.stats <strong>import </strong>wasserstein_distance</pre><pre><strong>for </strong>xi <strong>in </strong>[1e-1,1e-2,1e-3]:<br>    source, target = generate_source_target(xi)<br>    wass_1d = wasserstein_distance(source,target)<br>    print(round(np.abs(wass_1d-xi),2))</pre><pre>0.0<br>0.0<br>0.0</pre><p>To summarize, we now have a source domain for which we can learn a perfect classifier and a target domain that can be made arbitrary close to it and for which there exists a perfect classifier too. Now, the question we ask is: Can transfer learning succeed in this seemingly very favorable case?! Well, as you may have already guessed, no, it can’t, and the reason for this hides in the non-estimable term that I have mentioned before. Indeed, in this setup, a classifier that would be good for both domains <strong>simultaneously </strong>does not exist<strong>, </strong>whatever you may do. Actually, its lowest possible error will be exactly equal to 1-ξ which, once again, can be made arbitrary close to 1 by manipulating ξ accordingly. As put by the authors of the paper establishing the foundation of the transfer learning theory,</p><blockquote>“When there is <strong>no </strong>classifier that performs well on <strong>both</strong> the source and target domains, we <strong>cannot </strong>hope to <strong>find a good target model</strong> by training <strong>only </strong>on the source domain.”</blockquote><p>Sadly enough, this<strong> </strong>is true for<strong> all</strong> transfer learning approaches that do not have access to labels in the target domain.</p><h4>After-note</h4><p>While this article may seem a little pessimistic, its main goal, however, is to provide a general understanding of transfer learning theory to the interested reader and not to deceive him or her from using transfer learning methods. Indeed, transfer learning algorithms are usually quite efficient in practice as they are often applied to datasets that have a strong semantic connection between them. In this case, assuming that the non-estimable term is small is reasonable as, most likely, there exists a good classifier that will work well for both domains. However, as it often happens in life, you will never know if this is the case unless you try it, and that’s what makes the beauty of it.</p><p><strong>P.S.</strong> Let me know in the comments if you want a more in-depth discussion of the technical details behind any paper on transfer learning theory that you may be interested in.</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=27dcb8095670" width="1" height="1" alt=""><hr><p><a href="https://medium.com/data-science/why-transfer-learning-works-or-fails-27dcb8095670">Why transfer learning works or fails?</a> was originally published in <a href="https://medium.com/data-science">TDS Archive</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
    </channel>
</rss>