WebsiteContactHarvester

Crawl websites for contact information. Extract email, phone, facebook, twitter.

How to use

Clone the repo.
Restore NPM packages.
Update the sites to crawl in the sitesToCrawl.js file.
Execute node app.js
Harvested contact info will be placed into the ./output directory.

Output

Currently all potential phone numbers, email (mailto) address, twitter, and facebook URLs are harvested from retrieved HTML files. You can harvest additional data by modifying the harvestContactInfo method of the websiteContactHarvester.js class. The harvested data is saved to the ./output directory, one .json file per domain in the source sitesToCrawl.js file. You can use a tool like https://konklone.io/json/ to convert the .json files into .csv files.

Roadmap

Produce a .csv output file in addition to the .json files.
Eliminate duplicate values from the output files.

Name		Name	Last commit message	Last commit date
Latest commit History 6 Commits
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
app.js		app.js
package.json		package.json
sitesToCrawl.js		sitesToCrawl.js
websiteContactHarvester.js		websiteContactHarvester.js

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

WebsiteContactHarvester

How to use

Output

Roadmap

About

Releases

Packages

Languages

License

aaronhoffman/WebsiteContactHarvester

Folders and files

Latest commit

History

Repository files navigation

WebsiteContactHarvester

How to use

Output

Roadmap

About

Topics

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages