Live Online Course – 2nd Edition
Data Mining in R: Web Scraping and Text Analysis
January 9th, 16th, 23rd, and 30th, 2026
Are you interested in this course?
Please fill in the form below, and we will let you know as soon as the next edition is announced.
You can also use this form to request this as an in-house course for your group.
Click here to express interest in this course
Researchers often require access to large datasets from various sources, such as biodiversity databases, environmental monitoring websites, or online repositories. Manually collecting such data is time-intensive and inefficient. Web scraping can automate the extraction of scientific data from public databases, scientific publications, and organisational websites. For instance, one can scrape weather data or geological survey results to analyse trends or share findings with collaborators.
Establishing networks and promoting lifelong scientific learning often require analysing large volumes of text, such as conference abstracts, published papers, or grant announcements, to identify trends, common research interests, or potential collaborators. Text mining and Natural Language Processing (NLP) techniques can process large text datasets to identify key topics or perform Sentiment Analysis to assess public or academic opinions on certain scientific issues.
This course leverages the power of the R programming language, known for its extensive library of over 20,000 packages for statistical analysis, visualisation, text analysis, and machine learning. It focuses on teaching participants how to collect, process, and analyse web data effectively.
Key Highlights:
This course includes a range of activities such as web-scraping demos, live-coding sessions, interactive quizzes, and practical exercises to work individually or in a group. Active participation and contribution are recommended.
In this course we highlight the importance of prioritising the use of official APIs first. Web scraping is covered only when no suitable or dedicated API exists and more specifically for scientific, research and educational purposes. We also cover the areas of responsible web scraping such as respecting terms of service, copyright and server overload.
Places are limited to 15 participants.
Part 1: Web Scraping with R
Part 2: Text Analysis and NLP
Course participants are expected to have working knowledge of the R programming language. Prior experience in basic data analysis (such as data manipulation and visualisation) would help the learning experience but is not required
Participants should have their own computer with R, RStudio and the relevant packages installed. Instructions for the technical setup will be circulated before the course.
Online live sessions on January 9th, 16th, 23rd, and 30th, 2026.
From 9:30 to 13:00 (Madrid time zone).
Total course hours: 22
14 hours of online live lessons, plus 8 hours of independent work on exercises.
This course is equivalent to 1 ECTS (European Credit Transfer System) at the Life Science Zurich Graduate School. The recognition of ECTS by other institutions depends on each university or school.
In the live sessions we will combine online lectures with hands-on computational exercises in R.
Live sessions will be recorded. However, attendance to the live sessions is required to obtain the course certificate.
This course will be delivered in English.
Dr. Nicolas Attalides
Freelance
UK


