Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

[D] Machine Learning Approach in detecting if companies are the same

I have a very large dataset of shipment data where the company names are not normalized (e.g. companies that are supposed to be the same are treated different, like Walmart Inc., Walmart Incorporated, Wallmart, WalmartInc.). A simple string normalization like regex would not do good on this.

I have thought of TEXT SIMILARITY approach (Levenshtein Distance, Waro-Jinkler, etc.) which theoretically would work but would not do good in practice. One is that you should set a threshold and thresholds are different for each of them and various problems would arise.

Example:

  1. Large and Short Company Names would skew the threshold: (Walmart Inc – Walmart Inc. vs ABC Co. – In behalf of ABC Group of Co.)
  2. Almost similar company names that are supposed to be different (ABC Company Thailand vs ABC Company Taiwan)

The problem for #1 is that Thresholding for text similarity ratio is tricky.

The problem for #2 is that these companies have high ratio but are supposed to be different companies shipping different products (for example, ABC Thailand ships dresses while ABC Taiwan ships gadgets).

I have shipping data that looks like this

company name products company postal address country zip code
ABC Company Thailand 1x dress pink Bangkok Thailand Thailand 11100
ABC Company Taiwan 20x Phones Taipei, Taiwan Taiwan 00291
Walmart California Inc. 100kgs banana California California 9929
In behalf of Walmart CaliforniaInc 200kgs meat California California 9929

I am thinking of a solution that uses TEXT similarity metrics but across fields that could indicate that they are the same company (such as country, zip code, even products).

My proposed solution is

– a new entry is compared to a constructed table consisting of columns that are distinguishing features (company name, zip code, country for example)

– the new entry is only compared using the company name. the highest similarity is returned. And text similarity across different columns on new entry and selected data is produced.

– text similarity ratio/points of these two is fed to a classifier that tells if they are similar companies or not. Basically, the input for the classifier is the text similarity ratio of the new entry and the nearest company name from the list.

Any easier approach? The approach should be able to tackle both an existing large data and new entry (for example, deduplication does not seem to tackle addition of new entries). Thanks!

submitted by /u/sarmientoj24
[link] [comments]

[P] Sliding batch of training data to look at past X number of values

Hi all. I am very new to the ML world but I have been thinking about mapping a project out using a bunch of my own blood glucose data I have and did not know where to start with a sliding batch of input data? In order for a model to predict the next few values, it would need to know what happen in the last X number of minutes/ X number of data points. Does anyone have any links or textbooks to checkout on models like this? Anything is appreciated!

submitted by /u/dabirdman360
[link] [comments]

[P] Deep learning in Brancher.

Brancher (brancher.org) is a new framework for deep probabilistic inferece based on PyTorch.

In this new tutorial, we show how to use Brancher to build stochastic deep learning models

https://colab.research.google.com/drive/1YNwZpJgrsicK3Pz8igAktbdocm-gA9mV

We hope you’ll enjoy it! Let us know if you have questions or suggestions for future improvements either here or on Twitter: @pybrancher

submitted by /u/LucaAmbrogioni
[link] [comments]

[D] Any references on building models that infer real world physics?

Hi! I have been wondering about how to build a model that learns to infer the physics of environments from its input.

More specifically: how to build a model that takes input videos of, let’s say, objects falling down ( a glass being pushed from a table, an apple falling from a tree, etc.) and then makes the inference that things tend to fall to the ground? The model does not have to figure out any mathematical formula for calculating the gravitational force. Learning the common sense physics as humans do is what I am primarily interested in.

Has anyone come across research that addresses this problem or a similar problem? I would greatly appreciate any help. Thanks in advance! 🙂

submitted by /u/darthmeshkat
[link] [comments]

[D] Generative Adversarial Networks – The Story So Far

Hi everyone. I just published a new blog post which talks about the evolution of GANs over the last few years. You can check it out here.

I think it’s fascinating to see sample images generated from these models side by side. It really does give a sense of how fast this field has progressed. In just five years, we’ve gone from blurry, grayscale pixel arrays that vaguely resemble human faces to thispersondoesnotexist, which can easily fool most people on first glance.

Apart from image samples, I’ve also included links to papers, code, and other learning resources for each model. So this article could be an excellent place to start if you’re a beginner looking to catch up with the latest GAN research.

I hope you enjoy it!

submitted by /u/iyaja
[link] [comments]

[P] Multilabel Classifier With Closely Related Labels

Hey, I’m an industry outsider/hobbyist and trying to use AutoML text as a ternary classifier to prioritize incoming service request. The thing I realized is that multi-label classification appears to be for unrelated labels, but I’m dealing with labels that sit on a line, “low priority”, “medium priority”, “high priority”.

I don’t think this is the same problem as classifying with categorical labels such as “car”, “boat”, “plane”. The distance between low and high priority is much greater than medium, but the distance between a car, boat and plane are likely arbitrary. Is there any way to capture this? The only thing I’ve thought of so far is using a binary decision tree to check if low priority or not. If not, then check high priority. If not, then it’s assumed to be medium priority. Or does even that not work?

Sorry, I’m still trying to learn terminology and more on this subject. I’m not even sure what the name of the property I’m referring to is called.

submitted by /u/MaxxBreak
[link] [comments]

Another triple for the DeepRacer League brings more world records and the first female winner!

The AWS DeepRacer League is the world’s first global autonomous racing league, open to anyone. Developers of all skill levels can compete in person at 22 AWS events globally, or online via the AWS DeepRacer console, for a chance to win an expense paid trip to re:Invent 2019, where they will race to win the Championship Cup 2019.

Last week the AWS DeepRacer League visited three cities around the world – Washington D.C, USA, Taipei, Taiwan, and Tokyo, Japan. Each race spanned multiple days, providing developers with numerous opportunities to record a winning lap time.

The first female winner and another world record

The Tokyo race was the biggest one yet. Over 20,000 AWS customers came to the AWS Summit at the Makuhari Messe, located just outside of the city, for three days of learning, hands-on labs, and networking. There were two DeepRacer tracks for developers to race on throughout the summit, virtual racing pods, and multiple workshops to learn how to build a DeepRacer model.

Virtual racing pods, for customers to build models and learn more about the AWS DeepRacer league.

Hundreds of developers tested out their model on the tracks, but none could take the top spot from our first female winner, sola@DNP, who took home the cup with a world record winning time of 7.44 seconds – that means DeepRacer is travelling at the equivalent of roughly 100mph if scaled up to a real size car! Here is sola@DNP celebrating on the podium with her teammates. Check out the lightning fast winning lap!

She came to the AWS Summit as part of a team created at her company DNP (Dai Nippon Printing, a Japanese printing company operating in areas such as Information Communications, Lifestyles and Industrial Supplies and, Electronics). 28 of them placed on the leaderboard, with the top 3 all being from the team – 2 of them beating the previous world record (7.62 seconds) set just the week before at Amazon re:MARS.

To prepare for such a strong showing, DNP created DeepRacer study groups where employees share their knowledge and newly acquired machine learning skills. They see DeepRacer as a fun and engaging way to grow their engineers’ skills in AI.

“We currently have around 2000 IT personnel in the group and fewer than 200 employees experienced in working with AI. We want to double the number within 5 years.” – Mr. Shinichiro Fukuda, Deputy Director, C & I Center, DNP Information Innovation Division

Source: https://japan.zdnet.com/article/35137517/

Race Stats

The race in Japan was the most competitive yet. The top 33 competitors achieved lap times of under 10 seconds, the top 17 were under 9 seconds and the top 4 were under 8 seconds – breaking the world record twice! Check out the fast times and final results from the race on the Tokyo Leaderboard.

Washington DC

The AWS Public Sector Summit in Washington DC on June 10 also had an exciting race and there was a familiar face back on the tracks – our second place winner at Amazon re:MARS, John Amos. John narrowly missed out on the win and the opportunity to compete at re:Invent, in Las Vegas, when Anthony Navarro beat his world record time in the last few minutes of racing. In Washington, he took the lead early on and held his position steadfast in his pursuit of the win and will now be winging his way to re:Invent with the other Summit and Virtual circuit winners. He is really enjoying his AWS DeepRacer experience and has a new hobby to boot!

“I think everyone should have a hobby and this is a healthy one. There’s lots of stuff you can get addicted to, but with this you’re out there running models using technology. I’m used to playing video games online, but this helps it become real. What you do in the reward function impacts what’s happening with the car, so taking it from the simulator out onto the track is just exhilarating, and who doesn’t love a good challenge?”

Taipei

In Taipei, developers were also burning rubber and posting fast times on the leaderboard, and the winner took the top spot by a narrow margin (just 0.04 of a second), and in an exciting last few minutes of the race! He was Roger@NCTU_CGI, with a winning time of 8.734 seconds. Congratulations to all of the winners this week, it’s going to be an exciting final round at re:Invent 2019.

The AWS DeepRacer League Summit Circuit is in the homestretch

The AWS DeepRacer League Summit circuit only has five more races, (Hong Kong, Cape Town, New York, Mexico City, and Toronto) before the finale in Las Vegas, and it is shaping up to be an exciting event. Join the league at one of the 5 remaining races on the summit circuit, or race online in the virtual circuit today for your chance to win your trip to compete!


About the Author

Alexandra Bush is a Senior Product Marketing Manager for AWS AI. She is passionate about how technology impacts the world around us and enjoys being able to help make it accessible to all. Out of the office she loves to run, travel and stay active in the outdoors with family and friends.