Gypsylist

gypsylist.py is a web scraper for nomadlist.com, made to avoid website restrictions.

nomadlist.com is a website with a lot of information for digital nomad people, to find the best places to live and work remotely as a location independent remote worker. Unfortunately most of these contents are restricted if you are not member of this website.

This script doesn't cover all of the information retrievable from the website, but it's just an entry point to evaluate this without to sign up.

Installation

Before to use gypsylist you have to install some requirements:

pip3 install -r requirements.txt

Additionally, having selenium as dependency, you have also to setup the browser driver. To install this, please, take a look here: https://www.selenium.dev/documentation/webdriver/getting_started/install_drivers/.

Now you should be ready to run the script.

Usage

To use gypsylist, at first, browse the nomadlist.com website and apply the filters you need to do your research. Now, get the url path from the address bar of your browser (as shown below):

And use this to scrape with gypsylist:

./gypsylist.py --path "safe-places-for-remote-workers-to-live?sort=cost_for_nomad_in_usd&order=asc" --emoji

This is going to be the expected result:

#1
🏙️  city: Lisbon
🌎 country: Portugal
⭐️ overall: 4/5
💵 cost: 4/5
📡 internet: 5/5
😀 fun: 5/5
👮 safety: 4/5

...

#440
🏙️  city: Zurich
🌎 country: Switzerland
⭐️ overall: 3/5
💵 cost: 1/5
📡 internet: 5/5
😀 fun: 4/5
👮 safety: 4/5

#441
🏙️  city: Leiden
🌎 country: Netherlands
⭐️ overall: 3/5
💵 cost: 1/5
📡 internet: 5/5
😀 fun: 4/5
👮 safety: 4/5

#442
🏙️  city: Honolulu, Hawaii
🌎 country: United States
⭐️ overall: 4/5
💵 cost: 1/5
📡 internet: 5/5
😀 fun: 5/5
👮 safety: 4/5

#443
🏙️  city: Lake Tahoe, CA
🌎 country: United States
⭐️ overall: 3/5
💵 cost: 1/5
📡 internet: 5/5
😀 fun: 4/5
👮 safety: 4/5

(Always remember --emoji). Have fun!

Known Issues

This is not what you can call "a well written code" (sorry Gods of programming for this). For this reason there are several code smell or bugs that are not under review (due to the short time I dedicated to write the script).

Using --headless / -H parameter to set the browser in headless mode, you will retrieve just the first page contents from the website.

A web scraper for nomadlist.com, made to avoid website restrictions.

Related tags

Overview

Gypsylist

Installation

Usage

Known Issues

Owner

Alessio Greggi

Scrap the 42 Intranet's elearning videos in a single click

A modern CSS selector implementation for BeautifulSoup

EBay-email-tracker - Scapes an entire search page of a particular item on eBay and sends regular updates to an email address

爬虫案例合集。包括但不限于《淘宝、京东、天猫、豆瓣、抖音、快手、微博、微信、阿里、头条、pdd、优酷、爱奇艺、携程、12306、58、搜狐、百度指数、维普万方、Zlibraty、Oalib、小说、招标网、采购网、小红书》

Automated data scraper for Thailand COVID-19 data

Basic-html-scraper - A complete how to of web scraping with Python for beginners

Generate a repository with mirror links for DriveDroid app

Example of scraping a paginated API endpoint and dumping the data into a DB

Newsscraper - A simple Python 3 module to get crypto or news articles and their content from various RSS feeds.

OSTA web scraper, for checking the status of school buses in Ottawa

Amazon scraper using scrapy, a python framework for crawling websites.

FilmMikirAPI - A simple rest-api which is used for scrapping on the Kincir website using the Python and Flask package

A multithreaded tool for searching and downloading images from popular search engines. It is straightforward to set up and run!

Console application for downloading images from Reddit in Python

Using Python and Pushshift.io to Track stocks on the WallStreetBets subreddit

A python script to extract answers to any question on Quora (Quora+ included)

New World Market Scraper

A Simple Web Scraper made to Extract Download Links from Todaytvseries2.com

An automated, headless YouTube Watcher and Scraper

Parsel lets you extract data from XML/HTML documents using XPath or CSS selectors