Crawl websites for contact information. Extract email, phone, facebook, twitter.
- Clone the repo.
- Restore NPM packages.
- Update the sites to crawl in the
sitesToCrawl.js
file. - Execute
node app.js
- Harvested contact info will be placed into the
./output
directory.
Currently all potential phone numbers, email (mailto) address, twitter, and facebook URLs are harvested from retrieved HTML files.
You can harvest additional data by modifying the harvestContactInfo
method of the websiteContactHarvester.js
class.
The harvested data is saved to the ./output
directory, one .json file per domain in the source sitesToCrawl.js
file.
You can use a tool like https://konklone.io/json/ to convert the .json files into .csv files.
- Produce a .csv output file in addition to the .json files.
- Eliminate duplicate values from the output files.