Skip to main content

Blog

Learn About Our Meetup

5000+ Members

MEETUPS

LEARN, CONNECT, SHARE

Join our meetup, learn, connect, share, and get to know your Toronto AI community. 

JOB POSTINGS

INDEED POSTINGS

Browse through the latest deep learning, ai, machine learning postings from Indeed for the GTA.

CONTACT

CONNECT WITH US

Are you looking to sponsor space, be a speaker, or volunteer, feel free to give us a shout.

Author: torontoai

[D] Reinforcement learning, fast and slow

Hi everyone. I work on the neuroscience team at DeepMind. We’ve just published a new paper “Reinforcement learning, fast and slow” that reviews new techniques in deep reinforcement learning aiming to close the gap in the learning speed between humans and AI. Specifically, we look at how approaches like episodic deep RL and meta-reinforcement learning could unlock greater understanding in psychology and neuroscience by investigating the connection between fast and slow forms of deep RL.

If you’re interested in how AI and neuroscience can intersect, you can read the full paper here (available open access) – let us know what you think!

submitted by /u/JaneXWang
[link] [comments]

[D] How would one detect data leakage in someone else’s model?

Pure hypothetical. Let’s say I have someone’s model (i.e. their final model weights) and also their train and test set. I don’t have any additional validation readily available.

What kind of heuristics can be used to evaluate if there was data leakage from the test set?

I’d like to distinguish the two cases 1) there is data leakage, 2) the model is really good. Based on just performance metrics on the test/train set, I dont feel like I can distinguish these two cases. Would it be impossible to tell without additional validation data?

submitted by /u/CrazyAsparagus
[link] [comments]

[D] Deep Network Architectures

I’ve recently started looking at Deep Learning methods (specifically LSTM) for doing some time series analysis. While I understand the concepts involved in the network in isolation, one thing that confuses me is that how do we derive the architectures that work for a task.

I did some research into this and an answer that usually comes up is to dig into the literature and see what architect similar experiments used and start from there. My question is that is there a qualitative metric or a method that guides us for creating optimal architectures without using the layers as Lego bricks and connecting them hoping to get the best results.

submitted by /u/kalakesri
[link] [comments]

[D] lMetrics for similar classes classification

Hello,

I am having trouble finding right metrics for my problem.

I am predicting tags from text (multiclass classification) and I need to come up with metric that allows me to evaluate models.

I am taking top n classes with biggest probability as model output. I can’t monitor just precision/recall/f1 because I have similar classes. For example lets say that text X have classes [ocean, fish, boat] and my predictions are [sea, shark, boat]. I would say that this is correctly classified because sea and ocean are very similar, same for shark and fish. I am currently using word2vec and taking mean cos similarity between each prediction and label which gives maximum cos with given prediction. The problem with this metric is (I suspect) that its value decrease as I increase n (top predictions).

I think that good metrics would be something that combine cos metrics with precision.

Thanks

submitted by /u/matej1408
[link] [comments]

Goodwill Farming: Startup Harvests AI to Reduce Herbicides

Jorge Heraud is an anomaly for a founder whose startup was recently acquired by a corporate giant: Instead of counting days to reap earn-outs, he’s sowing the company’s goodwill message.

That might have something to do with the mission. Blue River Technology, acquired by John Deere more than a year ago for $300 million, aims to reduce herbicide use in farms.

The effort has been a calling to like-minded talent in Silicon Valley who want to apply their technology know-how to more meaningful problems than the next hot app, said Heraud, who continues to serve as Blue River’s CEO.

“We’re using machine learning to make a positive impact on the world. We don’t see it as just a way of making a profit. It’s about solving problems that are worthy of solving — that attracts people to us,” he said.

Heraud and co-founder Lee Redden, who continues to serve as Blue River’s CTO, were attending Stanford University in 2011 when they decided to form the startup. Redden was pursuing graduate studies in computer vision and machine learning applied to robotics while Heraud was getting an executive MBA.

The duo’s work formed one of the early success stories of many for harnessing NVIDIA GPUs and computer vision to tackle complex industrial problems with big benefits to humanity.

“Growing food is one of the biggest and oldest industries — it doesn’t get bigger than that,” said Ryan Kottenstette, who invested in Blue River at Khosla Ventures.

Herbicide Spraying 2.0

As part of tractor giant John Deere, Blue River remains committed to herbicide reduction. The company is engaged in multiple pilots of its See & Spray smart agriculture technology.

Pulled behind tractors, its See & Spray machine is about 40 feet wide and covers 12 rows of crops. It has 30 mounted cameras to capture photos of plants every 50 milliseconds and process them through its on-board 25 Jetson AGX Xavier supercomputing modules.

As a tractor pulls at about 7 miles per hour, according to Blue River, the Jetson Xavier modules running Blue River’s image recognition algorithms need to decide whether images fed from the 30 cameras are a weed or crop plant quicker than the blink of an eye. That allows enough time for the See & Spray’s robotic sprayer — it features 200 precision sprayers — to zap each weed individually with herbicide.

“We use Jetson to run inference on our machine learning algorithms and to decide on the fly if a plant is a crop or a weed, and spray only the weeds,” Heraud said.

GPUs Fertilize AgTech

Blue River has trained its convolutional neural networks on more than a million images and its See & Spray pilot machines keep feeding new data as they get used.

Capturing as many possible varieties of weeds in different stages of growth is critical to training the neural nets, which are processed on a “server closet full of GPUs” as well as on hundreds of GPUs at AWS, said Heraud.

Using cloud GPU instances, Blue River has been able to train networks much faster. “We have been able to solve hard problems and train in minutes instead of hours. It’s pretty cool what new possibilities are coming out,” he said.

Among them, Jetson Xavier’s compact design has enabled Blue River to move away from using PCs equipped with GPUs on board tractors. John Deere has ruggedized the Jetson Xavier modules, which offer some protection from the heat and dust of farms.

Business and Environment

Herbicides are expensive. A farmer spending a quarter-million dollars a year on herbicides was able to reduce that expense by 80 percent, Heraud said.

Blue River’s See & Spray can take the place of conventional, or aerial spraying of herbicides, which blankets entire crops with chemicals, something most countries are trying to reduce.

See & Spray can reduce the world’s herbicide use by roughly 2.5 billion pounds, an 80 percent reduction, which could have huge environmental benefits.

“It’s a tremendous reduction in the amount of chemicals. I think it’s very aligned with what customers want,” said Heraud.

 

Image credit: Blue River

The post Goodwill Farming: Startup Harvests AI to Reduce Herbicides appeared first on The Official NVIDIA Blog.

[P] Looking for Kaggle/Project Buddies

Hey all,

Not sure if this is the best place to post this, but I am looking to enhance my portfolio and ML skills more importantly through Kaggle competitions and side projects.

About me:

recent cornell grad with coursework experience in ML, did an internship dealing with NLP and CV applications.

currently following along with cs231 and 230 to improve my deep learning skills.

As I mentioned, looking for like minded individuals to partake in side projects or kaggle contests with!

submitted by /u/csjobseeker1
[link] [comments]

[Discussion] Problem: Classify variables in large data set

Hi everyone. I have been presented with a problem and I’m hoping someone here could provide some advice or a direction in which to look.

The problem is this: Given a large data set with, say, 10,000 columns (so 10,000 variables or factors), classify some variables as type A and the others as type B.

More specifically, the data set contains customer data, and some of the variables are personal information such as customer address, SSN, etc. and need to be classified as private. Since there are so many variables, one cannot simply identify the private ones and mark them as private/not private. The process needs to be automated, and we have training sets that are already classified that could be used for training a model to recognize private variables in future, unmarked datasets.

My problem is that I do not know what machine learning techniques are appropriate for this type of classification task. My understanding is that typically classification methods will classify the record value (row variable, in this case a customer) according to the values of the variable. My problem seems to be the inverse.

Furthermore, it would be interesting to be able to classify each private variable that appears in the future data set. For example, if one column contains SSNs, can we identify it and mark the column correctly?

Thank you in advance for any comments or advice.

submitted by /u/statsaccount
[link] [comments]