This corpus contains 492,993 aligned sentences extracted by pairing Simple English Wikipedia with English Wikipedia. These source data were downloaded in May 2016.
The form of each line in the corpus:
original sentence <TAB> simple sentence <TAB> similarity score
For questions, please contact Tomoyuki Kajiwara at Tokyo Metropolitan University.