AI Research Scientist – CMOR group corporation – Toronto, ON
From Indeed – Thu, 04 Apr 2019 18:36:14 GMT – View all Toronto, ON jobs
In a traditional human rights investigation, researchers travel to a region, conduct interviews, visit crime scenes, examine court records, and collect hospital or autopsy records.
While that painstaking approach still constitutes a major part of Human Rights Watch’s work, the U.S.-based nonprofit is also exploring new technological methods — including AI — for its investigations, said Fred Abrahams, an associate director.
“It would be irresponsible of us not to do that,” Abrahams said in a talk attended by more than 100 people at last month’s GPU Technology Conference. “We must explore every opportunity we can to get the goods to report on these human rights violations.”
These new tools include remote sensing via satellite and drone data, analytics from public datasets, and investigations using videos and photos posted to social media. Remote sensing is essential in situations where researchers can’t access a conflict zone or closed country — a major issue for the human rights and humanitarian community.
“We can’t document it if we can’t get there,” said Josh Lyons, director of geospatial analysis at Human Rights Watch. “If the people are in hiding or they’re dead, there’s no way to document that case.”
To push this work forward, the nonprofit is partnering with Element AI, a global AI software provider cofounded in 2016 by deep learning pioneer Yoshua Bengio. The company has a team in London focused on building AI for social good.
In addition to using NVIDIA GPUs in Element AI’s data center, Human Rights Watch is using two NVIDIA DGX Stations, provided in 2018 by NVIDIA, to further their efforts.
“The hardware will allow us to make it work,” Abrahams said.
There are hundreds of satellites orbiting and observing the Earth. Aerial imagery can show geographic features, human settlements and forces like flood and fire. Comparing how a region looks at one moment in time compared to another can be critical for human rights investigations — but the influx of data is too vast for any individual to go through.
At GTC, Lyons shared how Human Rights Watch was able to use thermal data from environmental satellites to begin monitoring the outbreak of ethnic violence in Myanmar in 2017, just hours after the first reports of conflict. Combined with aerial images, the organization was able to detect a pattern of burned Rohingya villages across the region.
This digital evidence helped on-the-ground researchers corroborate the testimony of the Muslim minority community targeted by the authorities. By pinpointing the exact date and time that a village began burning, investigators could better quantify the scale of violence and begin to determine who the perpetrators were.
But it takes an expert eye — or a neural network — to tell the difference between smoke plumes and puffy white clouds.
“Most of the time, it’s my eyes that are doing the analysis,” Lyons said. “The DGX immediately gives us the ability to scale.”
A deployed deep learning model that analyzes satellite or social media data could one day identify potential human rights abuses automatically from text and images and alert Human Rights Watch and humanitarian agencies.
However, though the proliferation of satellites and social media has led to a massive amount of new data for human rights investigators to parse, there’s still little labeled data to train neural networks. Looking at a satellite image of a smoke plume, “I know it’s a crime,” Lyons said. “But how do I tell the computer it’s a crime?”
That’s where Element AI’s expertise in deep learning can help. “By essentially cloning Josh’s visual cortex, we can have a huge impact,” said Julien Cornebise, director of research at Element AI. The company has also partnered with Amnesty International.
Cornebise and his team have also worked with Amnesty International on two projects: one to build neural networks to detect burned villages in Sudan, and another to parse Twitter data to study online abuse against women.
Human Rights Watch has been using the DGX Stations for photogrammetry, or converting 2D footage into 3D models, based on data collected from the nonprofit’s drones. The team is also developing and testing deep learning models to parse aerial imagery and social media data.
“We’re data rich and drowning in potential applications,” Lyons said. “The simple challenge is to prioritize.”
Potential uses include AI tools for processing archival footage dating back nearly 50 years, or making handwritten notes from Human Rights Watch investigators easier to translate or search.
These archives, particularly researchers’ notebooks, are “more or less locked in hard copy, paper form,” Lyons said. “Having such a system in place would be quite useful. It would give immeasurable value to future investigations.”
Having powerful deep learning systems onsite is also critical for Human Rights Watch to build AI tools analyzing sensitive datasets. For certain data such as forensic photographs or personal information, the organization is often not authorized to share the information with third parties — or host it on a remote server that falls under a specific geographic area or legal jurisprudence.
Lyons said, “The DGX Station hits that perfect sweet spot of being able to do large, robust data analysis in-house with sensitive data in a way that meets all of our legal and ethical privacy concerns.”
The above satellite image may look like clouds over a coastal community. However, an expert eye, or AI, can tell that the image shows smoke — revealing building fires in five villages in Myanmar’s Maungdaw township on the morning of September 15, 2017. Image courtesy of Human Rights Watch and Planet Labs Inc.
The post AI in the Sky Aids Feet on the Ground Spotting Human Rights Violations appeared first on The Official NVIDIA Blog.
Le 1er février 2018
L’Institut Vecteur tient à féliciter les membres de la première cohorte de son Programme de stagiaires diplômés. Ce programme lancé cette année se veut un catalyseur de collaboration entre les étudiants et les chercheurs spécialisés en apprentissage profond, en apprentissage automatique et, plus généralement, en intelligence artificielle. Il vient compléter d’autres programmes semblables qui verront le jour dans les établissements d’enseignement et les entreprises technologiques afin de favoriser un climat de collaboration propice au partage des idées et de l’expertise au sein de l’Institut Vecteur. La première cohorte est représentative des forces vives en présence dans les établissements et les universités de l’Ontario. On y trouve des étudiants à la maîtrise, des doctorants et des stagiaires postdoctoraux qui suivent de nouvelles pistes de recherche en apprentissage profond ou en apprentissage automatique.
Les candidats ont été évalués et sélectionnés en fonction de l’importance de leurs contributions antérieures à la recherche et de l’adéquation de leurs champs d’intérêt avec la vision, la mission et les capacités de recherche de Vecteur. Les étudiants retenus toucheront des honoraires en contrepartie de leur participation aux événements et aux activités de l’Institut. Devant le calibre élevé de cette première fournée, l’Institut espère enrichir et élargir le programme en 2019. La période de mise en candidature débutera vers la fin de l’année.
MEMBRES DE LA COHORTE 2018
● Ahmed Ashraf, Université de Toronto
● Alberto Camacho, Université de Toronto
● Andrew Boutros, Université de Toronto
● Aryan Arbabi, Université de Toronto
● Dmitry Marin, Université Western
● Elham Dolatabadi, Université de Toronto
● Ershad Banijamali, Université de Waterloo
● Ethan Jackson, Université Western
● Felix Berkenkamp, Université de Toronto
● Ga Wu, Université de Toronto
● Hassan Ashtiani, Université de Waterloo
● Ian Smith, Université de Toronto
● Jin Hee Kim, Université de Toronto
● Kathleen Houlahan, Université de Toronto
● Kiret Dhindsa, Université McMaster
● Laleh Soltan Ghoraie, Centre d’oncologie Princess Margaret
● Mahtab Ahmed, Université Western
● Matthew Tesfaldet, Université York
● Mehran Karimzadeh, Université de Toronto
● Mohamed Khairy Helwa, Université de Toronto
● Nabiha Asghar, Université de Waterloo
● Omar Boursalie, Université McMaster
● Petr Smirnov, Université de Toronto
● Rober Boshra, Université McMaster
● Robin Swanson, Université de Toronto
● Rodrigo Toro Icarte, Université de Toronto
● Sean Robertson, Université de Toronto
● Shazia Akbar, Université de Toronto
● Shehroz Khan, Université de Toronto
● SiQi Zhou, Université de Toronto
● Stavros Tsogkas, Université de Toronto
● Tristan Aumentado-Armstrong, Université de Toronto
https://vectorinstitute.ai/#partners
La Stratégie pancanadienne en matière d’IA dirigée par le CIFAR vise à encourager la collaboration entre l’Institut Vecteur, l’Amii (Alberta Machine Intelligence Institute) et Mila (Institut québécois d’intelligence artificielle).
Amazon Comprehend is a fully managed natural language processing (NLP) service that enables text analytics for important workloads. For example, analyzing market research reports for key market indicators or data that contains PII information. Customers that work with highly sensitive, encrypted data can now easily enable Comprehend to work with this encrypted data via an integration with the AWS Key Management Service.
AWS KMS makes it easy for you to create and manage keys and control the use of encryption across a wide range of AWS services and in your applications. AWS KMS is a secure and resilient service that uses FIPS 140-2 validated hardware security modules to protect your keys. AWS KMS is integrated with AWS CloudTrail to provide you with logs of all key usage to help meet your regulatory and compliance needs.
To enable Comprehend to use KMS keys to access data, the feature can be configured via the AWS Management console or the SDK and supports Amazon Comprehend asynchronous training and inference jobs. To get started you first need to create a key in the AWS KMS service. To learn more about how to create KMS keys, please visit: https://docs.aws.amazon.com/kms/latest/developerguide/create-keys.html
When you are configuring an asynchronous job, you can specify the KMS encryption key the Comprehend should use to access your data in S3. Below is an example of selecting a key with the alias “Comprehend” as part of configuring job details, in the Amazon Comprehend console:

To manage your AWS KMS keys, please visit the AWS KMS management portal or use the KMS SDK. For more information, please visit: AWS Key Management Service. To learn more about how to configure Comprehend jobs to work with KMS keys, please visit our documentation:
Nino Bice is a Sr. Product Manager leading product for Amazon Comprehend, AWS’s natural language processing service.
Recording video of memorable moments to share with friends and loved ones has become commonplace. But as anyone with a sizable video library can tell you, it’s a time consuming task to go through all that raw footage searching for the perfect clips to relive or share with family and friends. Google Photos makes this easier by automatically finding magical moments in your videos—like when your child blows out the candle or when your friend jumps into a pool—and creating animations from them that you can easily share with friends and family.
In “Rethinking the Faster R-CNN Architecture for Temporal Action Localization“, we address some of the challenges behind automating this task, which are due to the complexity of identifying and categorizing actions from a highly variable array of input data, by introducing an improved method to identify the exact location within a video where a given action occurs. Our temporal action localization network (TALNet) draws inspiration from advances in region-based object detection methods such as the Faster R-CNN network. TALNet enables identification of moments with large variation in duration, achieving state-of-the-art performance compared to other methods, allowing Google Photos to recommend the best part of a video for you to share with friends and family.
![]() |
| An example of the detected action “blowing out candles” |
Identifying Actions for Model Training
The first step in identifying magic moments in videos is to assemble a list of actions that people might wish to highlight. Some examples of actions include “blow out birthday candles”, “strike (bowling)”, “cat wags tail”, etc. We then crowdsourced the annotation of segments within a collection of public videos where these specific actions occurred, in order to create a large training dataset. We asked the raters to find and label all moments, accommodating videos that might have several moments. This final annotated dataset was then used to train our model so that it could identify the desired actions in new, unknown videos.
Comparison to Object Detection
The challenge of recognizing these actions belongs to the field of computer vision known as temporal action localization, which, like the more familiar object detection, falls under the umbrella of visual detection problems. Given a long, untrimmed video as input, temporal action localization aims to identify the start and end times, as well as the action label (like “blowing out candles”), for each action instance in the full video. While object detection aims to produce spatial bounding boxes around an object in a 2D image, temporal action localization aims to produce temporal segments including an action in a 1D sequence of video frames.
Our approach to TALNet is inspired by the faster R-CNN object detection framework for 2D images. So, to understand TALNet, it is useful to first understand faster R-CNN. The figure below demonstrates how the faster R-CNN architecture is used for object detection. The first step is to generate a set of object proposals, regions of the image that can be used for classification. To do this, an input image is first converted into a 2D feature map by a convolutional neural network (CNN). The region proposal network then generates bounding boxes around candidate objects. These boxes are generated at multiple scales in order to capture the large variability in objects’ sizes in natural images. With the object proposals now defined, the subjects in the bounding boxes are then classified by a deep neural network (DNN) into specific objects, such as “person”, “bike”, etc.
![]() |
| Faster R-CNN architecture for object detection |
Temporal Action Localization
Temporal action localization is accomplished in a fashion similar to that used by R-CNN. A sequence of input frames from a video are first converted into a sequence of 1D feature maps that encode scene context. This map is passed to a segment proposal network that generates candidate segments, each defined by start and end times. A DNN then applies the representations learned from the training dataset to classify the actions in the proposed video segments (e.g., “slam dunk”, “pass”, etc.). The actions identified in each segment are given weights according to their learned representations, with the top scoring moment selected to share with the user.
![]() |
| Architecture for temporal action localization |
Special Considerations for Temporal Action Localization
While temporal action localization can be viewed as the 1D counterpart of the object detection problem, care must be taken to address a number of issues unique to action localization. In particular, we address three specific issues in order to apply the Faster R-CNN approach to the action localization domain, and redesign the architecture to specifically address them.
TALNet in Action
As a consequence of these improvements, TALNet achieves state-of-the-art performance for both action proposal and action localization tasks on the THUMOS’14 detection benchmark and competitive performance on the ActivityNet challenge. Now, whenever people save videos to Google Photos, our model identifies these moments and creates animations to share. Here are a few examples shared by our initial testers.
![]() |
| An example of the detected action “sliding down a slide” |
![]() |
| An example of the detected actions “jump into the pool” (left), “twirl in a dress” (center) and “feed baby a spoonful” (right). |
Next steps
We are continuing work to improve the precision and recall of action localization using more data, features and models. Improvements in temporal action localization can drive progress on a large number of important topics ranging from video highlights, video summarization, search and more. We hope to continue improving the state-of-the-art in this domain and at the same time provide more ways for people to reminisce on their memories, big and small.
Acknowledgements
Special thanks Tim Novikoff and Yu-Wei Chao, as well as Bryan Seybold, Lily Kharevych, Siyu Gu, Tracy Gu, Tracy Utley, Yael Marzan, Jingyu Cui, Balakrishnan Varadarajan, Paul Natsev for their critical contributions to this project.