Shellcode_IA32: A Dataset for Automatic Shellcode Generation

A new dataset for shellcode generation from natural language comments is introduced, and standard neural machine translation methods are explored for baseline performance.

Open

Preview
Year: 2021
Venue: ACL (NLP4Prog) 2021 8
ArXiv: arxiv.org/abs/2104.13100
Authors: 6
Hosting: Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text: arxiv.org/abs/2104.13100v4ARXIV-DEFAULT
TL;DR: Semantic Scholar

Attribution policy →

Abstract

We take the first step to address the task of automatically generating shellcodes, i.e., small pieces of code used as a payload in the exploitation of a software vulnerability, starting from natural language comments. We assemble and release a novel dataset (Shellcode_IA32), consisting of challenging but common assembly instructions with their natural language descriptions. We experiment with standard methods in neural machine translation (NMT) to establish baseline performance levels on this task.

Authors

Domenico Cotroneo Pietro Liguori Erfan Al-Hossami Roberto Natella Bojan Cukic Samira Shaikh