{"id":494366,"date":"2018-07-10T06:51:52","date_gmt":"2018-07-10T13:51:52","guid":{"rendered":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/?post_type=msr-event&#038;p=494366"},"modified":"2025-08-06T11:57:00","modified_gmt":"2025-08-06T18:57:00","slug":"frontiers-in-ai-tal-ben-nun","status":"publish","type":"msr-event","link":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai-tal-ben-nun\/","title":{"rendered":"Frontiers in AI &#8211; Tal Ben-Nun"},"content":{"rendered":"\n\n<p>Auditorium<\/p>\n<p>21 Station Road<br \/>\nCambridge<br \/>\nCB1 2FB<\/p>\n<p>&nbsp;<\/p>\n<p>View the whole series on <a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"http:\/\/talks.cam.ac.uk\/show\/index\/64171\">talks.cam<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p>&nbsp;<span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n<p><span style=\"color: #ff6600\">Frontiers in Artificial Intelligence<\/span> is a series of public lectures at Microsoft Research Cambridge featuring leading researchers in the field, focusing on the cutting edge topics at the intersection of machine learning, statistics, and artificial intelligence. Students, scientists, and engineers in academia and industry are all welcome to join us for these exciting talks and the opportunity to socialize with the Cam-bridge AI\/ML community.<\/p>\n<p><strong><span style=\"color: #ff6600\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-568287 size-medium alignleft\" src=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-300x300.jpg\" alt=\"\" width=\"300\" height=\"300\" srcset=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-300x300.jpg 300w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-150x150.jpg 150w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-180x180.jpg 180w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-360x360.jpg 360w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004.jpg 500w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/span><\/strong><\/p>\n<p><strong><span style=\"color: #ff6600\">Neural Code Comprehension: A Learnable Representation of Code Semantics<\/span><\/strong><\/p>\n<p><strong><span style=\"color: #333300\">Tal Ben-Nun<\/span><\/strong><\/p>\n<p>In the era of \u201cBig Code\u201d, research is being conducted into automating the understanding of computer programs. Most of the current works base on techniques from Natural Language Processing and Deep Learning, which have been successful recently, attempting to process the code directly or using syntactic representations (e.g., ASTs and AST paths). However, to comprehend program semantics robustly, structural features of code have to be taken into account as well, including function calls, branching, and interchangeable order of statements. In this talk, I will present a novel processing technique to use Machine Learning for code semantics, and show how it applies to a variety of program analysis tasks. In particular, we stipulate that a robust distributional hypothesis of code applies to both human- and machine-generated programs. Following this hypothesis, we define an embedding space, inst2vec, based on an Intermediate Representation (IR) of the code that is independent of the source programming language. We provide a novel definition of contextual flow for this IR, leveraging both the underlying data- and control-flow of the program. We then analyze the embeddings quantitatively using analogies and clustering, and evaluate the learned representation on three different high-level tasks. We show that even without fine-tuning, a single Recurrent Neural Network (RNN) architecture and fixed inst2vec embeddings outperform specialized approaches for performance prediction (compute device mapping, optimal thread coarsening); and algorithm classification from raw code (104 classes), where we set a new state-of-the-art.<\/p>\n<h3><\/h3>\n<p><span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/116905\">Hoda Heidari &#8211; What can Fair ML learn from Economic Theories of Disruptive Justice?<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/108268\">Finale Doshi-Velez &#8211; Interpretability in Machine Learning: What it means, How we&#8217;re getting there.<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/97270\">Francis Bach\u00a0 &#8211; Optimal Algorithms for smooth and strongly convex distributed optimization in networks<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/80361\">Andreas Geiger &#8211; Probabilistic and Deep Models for 3D Reconstruction<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-ai-francesco-orabona\/\">Francesco Orabona &#8211; Coin Betting for Backprop without Learning rates and More<\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"http:\/\/talks.cam.ac.uk\/talk\/index\/73841\">Regina Barzilay &#8211; How Can NLP Help Cure Cancer?<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Aapo Hyvarinen &#8211; Nonlinear ICA using temporal structure: a principled framework for unsupervised deep learning\u00a0<\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Max Welling &#8211; Generalizing Convolutions for Deep Learning <\/a><\/p>\n<p>&nbsp;<span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Auditorium 21 Station Road Cambridge CB1 2FB &nbsp; View the whole series on talks.cam (opens in new tab) &nbsp;Opens in a new tab Frontiers in Artificial Intelligence is a series of public lectures at Microsoft Research Cambridge featuring leading researchers in the field, focusing on the cutting edge topics at the intersection of machine learning, [&hellip;]<\/p>\n","protected":false},"featured_media":0,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr_startdate":"2019-02-28","msr_enddate":"2019-02-28","msr_location":"Microsoft Research Cambridge","msr_expirationdate":"","msr_event_recording_link":"","msr_event_link":"","msr_event_link_redirect":false,"msr_event_time":"13:00","msr_hide_region":false,"msr_private_event":false,"msr_hide_image_in_river":0,"footnotes":""},"research-area":[13556],"msr-region":[],"msr-event-type":[197944],"msr-video-type":[],"msr-locale":[268875],"msr-program-audience":[],"msr-post-option":[],"msr-impact-theme":[],"class_list":["post-494366","msr-event","type-msr-event","status-publish","hentry","msr-research-area-artificial-intelligence","msr-event-type-hosted-by-microsoft","msr-locale-en_us"],"msr_about":"<!-- wp:msr\/event-details {\"title\":\"Frontiers in AI - Tal Ben-Nun\",\"backgroundColor\":\"grey\"} \/-->\n\n<!-- wp:msr\/content-tabs --><!-- wp:msr\/content-tab {\"title\":\"About\"} --><!-- wp:freeform --><p>Auditorium<\/p>\n<p>21 Station Road<br \/>\nCambridge<br \/>\nCB1 2FB<\/p>\n<p>&nbsp;<\/p>\n<p>View the whole series on <a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"http:\/\/talks.cam.ac.uk\/show\/index\/64171\">talks.cam<span class=\"sr-only\"> (opens in new tab)<\/span><\/a><\/p>\n<p>&nbsp;<span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n<p><span style=\"color: #ff6600\">Frontiers in Artificial Intelligence<\/span> is a series of public lectures at Microsoft Research Cambridge featuring leading researchers in the field, focusing on the cutting edge topics at the intersection of machine learning, statistics, and artificial intelligence. Students, scientists, and engineers in academia and industry are all welcome to join us for these exciting talks and the opportunity to socialize with the Cam-bridge AI\/ML community.<\/p>\n<p><strong><span style=\"color: #ff6600\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-568287 size-medium alignleft\" src=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-300x300.jpg\" alt=\"\" width=\"300\" height=\"300\" srcset=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-300x300.jpg 300w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-150x150.jpg 150w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-180x180.jpg 180w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-360x360.jpg 360w, https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004.jpg 500w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/span><\/strong><\/p>\n<p><strong><span style=\"color: #ff6600\">Neural Code Comprehension: A Learnable Representation of Code Semantics<\/span><\/strong><\/p>\n<p><strong><span style=\"color: #333300\">Tal Ben-Nun<\/span><\/strong><\/p>\n<p>In the era of \u201cBig Code\u201d, research is being conducted into automating the understanding of computer programs. Most of the current works base on techniques from Natural Language Processing and Deep Learning, which have been successful recently, attempting to process the code directly or using syntactic representations (e.g., ASTs and AST paths). However, to comprehend program semantics robustly, structural features of code have to be taken into account as well, including function calls, branching, and interchangeable order of statements. In this talk, I will present a novel processing technique to use Machine Learning for code semantics, and show how it applies to a variety of program analysis tasks. In particular, we stipulate that a robust distributional hypothesis of code applies to both human- and machine-generated programs. Following this hypothesis, we define an embedding space, inst2vec, based on an Intermediate Representation (IR) of the code that is independent of the source programming language. We provide a novel definition of contextual flow for this IR, leveraging both the underlying data- and control-flow of the program. We then analyze the embeddings quantitatively using analogies and clustering, and evaluate the learned representation on three different high-level tasks. We show that even without fine-tuning, a single Recurrent Neural Network (RNN) architecture and fixed inst2vec embeddings outperform specialized approaches for performance prediction (compute device mapping, optimal thread coarsening); and algorithm classification from raw code (104 classes), where we set a new state-of-the-art.<\/p>\n<h3><\/h3>\n<p><span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n<!-- \/wp:freeform --><!-- \/wp:msr\/content-tab --><!-- wp:msr\/content-tab {\"title\":\"Past Speakers\"} --><!-- wp:freeform --><p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/116905\">Hoda Heidari &#8211; What can Fair ML learn from Economic Theories of Disruptive Justice?<\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/108268\">Finale Doshi-Velez &#8211; Interpretability in Machine Learning: What it means, How we&#8217;re getting there.<\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/97270\">Francis Bach\u00a0 &#8211; Optimal Algorithms for smooth and strongly convex distributed optimization in networks<\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/80361\">Andreas Geiger &#8211; Probabilistic and Deep Models for 3D Reconstruction<\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-ai-francesco-orabona\/\">Francesco Orabona &#8211; Coin Betting for Backprop without Learning rates and More<\/a><\/p>\n<p><a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" target=\"_blank\" href=\"http:\/\/talks.cam.ac.uk\/talk\/index\/73841\">Regina Barzilay &#8211; How Can NLP Help Cure Cancer?<\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Aapo Hyvarinen &#8211; Nonlinear ICA using temporal structure: a principled framework for unsupervised deep learning\u00a0<\/a><\/p>\n<p><a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Max Welling &#8211; Generalizing Convolutions for Deep Learning <\/a><\/p>\n<p>&nbsp;<span id=\"label-external-link\" class=\"sr-only\" aria-hidden=\"true\">Opens in a new tab<\/span><\/p>\n<!-- \/wp:freeform --><!-- \/wp:msr\/content-tab --><!-- \/wp:msr\/content-tabs -->","tab-content":[{"id":0,"name":"About","content":"<span style=\"color: #ff6600\">Frontiers in Artificial Intelligence<\/span> is a series of public lectures at Microsoft Research Cambridge featuring leading researchers in the field, focusing on the cutting edge topics at the intersection of machine learning, statistics, and artificial intelligence. Students, scientists, and engineers in academia and industry are all welcome to join us for these exciting talks and the opportunity to socialize with the Cam-bridge AI\/ML community.\r\n\r\n<strong><span style=\"color: #ff6600\"><img class=\"wp-image-568287 size-medium alignleft\" src=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-content\/uploads\/2018\/07\/tal-004-300x300.jpg\" alt=\"\" width=\"300\" height=\"300\" \/><\/span><\/strong>\r\n\r\n<strong><span style=\"color: #ff6600\">Neural Code Comprehension: A Learnable Representation of Code Semantics<\/span><\/strong>\r\n\r\n<strong><span style=\"color: #333300\">Tal Ben-Nun<\/span><\/strong>\r\n\r\nIn the era of \u201cBig Code\u201d, research is being conducted into automating the understanding of computer programs. Most of the current works base on techniques from Natural Language Processing and Deep Learning, which have been successful recently, attempting to process the code directly or using syntactic representations (e.g., ASTs and AST paths). However, to comprehend program semantics robustly, structural features of code have to be taken into account as well, including function calls, branching, and interchangeable order of statements. In this talk, I will present a novel processing technique to use Machine Learning for code semantics, and show how it applies to a variety of program analysis tasks. In particular, we stipulate that a robust distributional hypothesis of code applies to both human- and machine-generated programs. Following this hypothesis, we define an embedding space, inst2vec, based on an Intermediate Representation (IR) of the code that is independent of the source programming language. We provide a novel definition of contextual flow for this IR, leveraging both the underlying data- and control-flow of the program. We then analyze the embeddings quantitatively using analogies and clustering, and evaluate the learned representation on three different high-level tasks. We show that even without fine-tuning, a single Recurrent Neural Network (RNN) architecture and fixed inst2vec embeddings outperform specialized approaches for performance prediction (compute device mapping, optimal thread coarsening); and algorithm classification from raw code (104 classes), where we set a new state-of-the-art.\r\n<h3><\/h3>"},{"id":1,"name":"Past Speakers","content":"<a href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/116905\">Hoda Heidari - What can Fair ML learn from Economic Theories of Disruptive Justice?<\/a>\r\n\r\n<a href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/108268\">Finale Doshi-Velez - Interpretability in Machine Learning: What it means, How we're getting there.<\/a>\r\n\r\n<a href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/97270\">Francis Bach\u00a0 - Optimal Algorithms for smooth and strongly convex distributed optimization in networks<\/a>\r\n\r\n<a href=\"https:\/\/talks.cam.ac.uk\/talk\/index\/80361\">Andreas Geiger - Probabilistic and Deep Models for 3D Reconstruction<\/a>\r\n\r\n<a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-ai-francesco-orabona\/\">Francesco Orabona - Coin Betting for Backprop without Learning rates and More<\/a>\r\n\r\n<a href=\"http:\/\/talks.cam.ac.uk\/talk\/index\/73841\">Regina Barzilay - How Can NLP Help Cure Cancer?<\/a>\r\n\r\n<a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Aapo Hyvarinen - Nonlinear ICA using temporal structure: a principled framework for unsupervised deep learning\u00a0<\/a>\r\n\r\n<a href=\"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/event\/frontiers-in-ai\/#\">Max Welling - Generalizing Convolutions for Deep Learning <\/a>\r\n\r\n&nbsp;"}],"msr_startdate":"2019-02-28","msr_enddate":"2019-02-28","msr_event_time":"13:00","msr_location":"Microsoft Research Cambridge","msr_event_link":"","msr_event_recording_link":"","msr_startdate_formatted":"February 28, 2019","msr_register_text":"Watch now","msr_cta_link":"","msr_cta_text":"","msr_cta_bi_name":"","featured_image_thumbnail":null,"event_excerpt":"Frontiers in Artificial Intelligence is a series of public lectures at Microsoft Research Cambridge featuring leading researchers in the field, focusing on the cutting edge topics at the intersection of machine learning, statistics, and artificial intelligence. Students, scientists, and engineers in academia and industry are all welcome to join us for these exciting talks and the opportunity to socialize with the Cam-bridge AI\/ML community. Neural Code Comprehension: A Learnable Representation of Code Semantics Tal Ben-Nun&hellip;","msr_research_lab":[199561],"related-researchers":[],"msr_impact_theme":[],"related-academic-programs":[],"related-groups":[],"related-projects":[],"related-opportunities":[],"related-publications":[],"related-videos":[],"related-posts":[],"_links":{"self":[{"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-event\/494366","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-event"}],"about":[{"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-event"}],"version-history":[{"count":6,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-event\/494366\/revisions"}],"predecessor-version":[{"id":1147092,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-event\/494366\/revisions\/1147092"}],"wp:attachment":[{"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/media?parent=494366"}],"wp:term":[{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=494366"},{"taxonomy":"msr-region","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-region?post=494366"},{"taxonomy":"msr-event-type","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-event-type?post=494366"},{"taxonomy":"msr-video-type","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-video-type?post=494366"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=494366"},{"taxonomy":"msr-program-audience","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-program-audience?post=494366"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=494366"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/cm-edgetun.pages.dev\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=494366"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}