Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
55 commits
Select commit Hold shift + click to select a range
d14c1e7
fix missing card(s) from douyin search
dale-wahl Nov 22, 2024
7c9eba1
update test for douyin
dale-wahl Nov 22, 2024
6adb0c9
douyin: skip embedded videos if ?modal_is in source URL (only display…
dale-wahl Nov 27, 2024
6372e0a
add douyin test
dale-wahl Nov 27, 2024
cfb261a
douyin: capture ?modal_id= video links!
dale-wahl Nov 28, 2024
402acb3
douyin catch "trending" video in search
dale-wahl Nov 29, 2024
889f6a8
Bump version
stijn-uva Nov 29, 2024
d9367ac
Skip live streams when capturing TikTok posts
stijn-uva Dec 12, 2024
888e60b
Bump version to 1.11.3
stijn-uva Dec 12, 2024
8269aed
Exclude TikTok live streams more thoroughly
stijn-uva Jan 15, 2025
b98d5ca
Pinterest module
stijn-uva Feb 4, 2025
68f8c78
Document code, extra check for complete data
stijn-uva Feb 5, 2025
dc3f866
Fix "extra check" for Pinterest data to not check toooo hard
stijn-uva Feb 6, 2025
920223b
Xiaohongshu/RedNote module
stijn-uva Feb 17, 2025
ef192bd
Use new generic traverse function elsewhere too
stijn-uva Feb 17, 2025
c2e8890
Xiaohongshu tests
stijn-uva Feb 17, 2025
08bfb84
Bump version
stijn-uva Feb 17, 2025
54ab065
Update readme with new platforms
stijn-uva Feb 18, 2025
2142f41
Enable modules for all pages where domain *ends with* module domain
stijn-uva Feb 19, 2025
6b0caa9
Improve code comments for RedNote module
stijn-uva Feb 19, 2025
1f8b478
Version +0.0.1
stijn-uva Feb 19, 2025
89980de
Fix Threads embed item capture
stijn-uva Feb 20, 2025
ac8be43
Fix #41
stijn-uva Feb 20, 2025
6fb25ec
Fix LinkedIn capture from initial search result page
stijn-uva Feb 20, 2025
37828ce
Delete test.js
stijn-uva Feb 20, 2025
6bf4c72
Fix inclusion of embedded Pinterest items
stijn-uva Feb 20, 2025
537ac67
Allow ranges for expected amount of items in tests
stijn-uva Feb 20, 2025
11551be
Update tests
stijn-uva Feb 20, 2025
4d1cedd
Version +0.0.1
stijn-uva Feb 20, 2025
1e1e895
RedNote comments module
stijn-uva Mar 12, 2025
2b12eeb
Bump version
stijn-uva Mar 12, 2025
56fe7e4
Don't run disabled modules on same domain as enabled
stijn-uva Mar 18, 2025
d586a13
Extra installation instructions
stijn-uva Mar 19, 2025
a529e24
Update threads domain
stijn-uva Apr 25, 2025
41ad74a
Bump version
stijn-uva Apr 25, 2025
3279e8a
Update link to guide
stijn-uva Jun 3, 2025
ccdddcf
tiktok: `duetInfo` no longer in objects
dale-wahl Sep 18, 2025
87fd4eb
tiktok: get JSON embed from rehydration; skip preload endpoint
dale-wahl Sep 18, 2025
43e711d
Version to 1.13.2
stijn-uva Sep 18, 2025
f6c7919
Add new data collection info to manifest
stijn-uva Sep 18, 2025
4e8a6ef
Fix Instagram capture from front page and single post page
stijn-uva Sep 28, 2025
0c40f32
Fix capture from Threads via embedded JSON
stijn-uva Sep 28, 2025
39651a2
Version++
stijn-uva Sep 28, 2025
6113fc3
Allow alternate Discover page URL for Douyin
stijn-uva Sep 30, 2025
b06ca88
version++
stijn-uva Sep 30, 2025
f3d7e65
Minimally improved upload instructions for 4CAT instance
dale-wahl Oct 3, 2025
472014f
Update issue templates
stijn-uva Oct 8, 2025
f4ad6f3
Update issue templates
stijn-uva Oct 8, 2025
79c4f76
Enhance README with data collection details
dale-wahl Nov 17, 2025
47d4cc8
skip non tweet results (e.g. timeline labels)
dale-wahl Nov 20, 2025
ec216ae
Merge branch 'master' into facebook-v2
uree Nov 27, 2025
e654c8f
no more duplicate saving of fb stories
uree Dec 2, 2025
4ed665e
include fb module
uree Dec 2, 2025
b10a83a
added fb comments module for downloading comments under posts
uree Jan 7, 2026
3d503c9
cleanup
uree Jan 7, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions .github/ISSUE_TEMPLATE/bug_report.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
---
name: Bug report
about: Something is not working or output seems wrong
title: ''
labels: ''
assignees: ''

---

Before submitting:
- Is this an issue with the output of 4CAT? Then please make an issue here instead: https://github.com/digitalmethodsinitiative/4cat/issues
- Is this an issue with the output of Zeehaven? Then please make an issue here instead: https://github.com/PublicDataLab/zeehaven/issues

*Describe your bug*
1. What version of Zeeschuimer are you using? You can find the version at the top of the interface, in the header.
2. What platform are your trying to capture data from?
3. Describe what should be collected, but is not (e.g. a specific post, or all posts from a specific user, or a particular type of item).
4. If possible, include a link to the page you are trying to capture data from.
2 changes: 1 addition & 1 deletion .zenodo.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
"license": "MPL-2.0",
"title": "Zeeschuimer",
"upload_type": "software",
"version": "v1.11.1",
"version": "v1.13.4",
"keywords": [
"scraping", "data capture", "4cat", "instagram", "tiktok"
],
Expand Down
36 changes: 28 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,9 @@ Currently, it supports the following platforms:
* [Imgur](https://imgur.com)
* [Douyin](https://douyin.com)
* [Gab](https://gab.com)
* [Truth Social](https://truth.social)
* [Pinterest](https://pinterest.com)
* [RedNote/Xiaohongshu](https://xiaohongshu.com)

Platform support requires regular maintenance to keep up with changes to the platforms. If something does not work, we
welcome issues and pull requests. See 'Limitations' below for some known limitations to data capture.
Expand All @@ -40,19 +43,29 @@ in any Firefox-based browser. If you want to run the latest development version
debugging console](https://www.youtube.com/watch?v=J7el77F1ckg) after cloning the repository locally.

## How to use
A [guide to using Zeeschuimer and 4CAT](https://tinyurl.com/nmrw-zeeschuimer-tiktok) is available. Basic instructions
A [guide to using Zeeschuimer and 4CAT](https://zeeschuimer.4cat.nl/) is available. Basic instructions
are as follows:

Install the browser extension in a Firefox browser. A button with the Zeeschuimer logo (a 'Z') will appear in the
browser toolbar. Click it to open the Zeeschuimer interface. Enable capturing for the sites you want to capture from.
Install the browser extension in a Firefox browser. A button with the Zeeschuimer logo (<img alt="Zeeschuimer's browser icon, a yellow 'Z' on a green background" src="images/zeeschuimer-16.png">) will appear in the browser toolbar. Click it
to open the Zeeschuimer interface. Enable capturing for the sites you want to capture from.

Next, simply browse a supported platform's site. You will see the amount of items detected per platform increase as you
browse. When you have the items you need, you can export the data as an [ndjson](https://ndjson.org) file, or upload it
to a 4CAT instance where a 4CAT dataset will be created from the uploaded items. You can then run 4CAT's analytical
Note that after installation in Firefox, the extension icon may not be immediately visible in the toolbar. If you can't
find Zeeschuimer's icon, look for the 'Extensions' icon (a puzzle piece); clicking it will show all available extensions
that are not shown in the main browser toolbar.

### Collecting data
Next, simply browse a supported platform's site. **Zeeschuimer collects data as your browser receives it**. It is
always best practices to refresh a page after you toggle collection on for a supported platform's site in the control
panel (any data loaded previously will not be collected until refreshed). You can force a refresh with Ctl + F5 on Windows
or Shift + Command + R on Mac. You will see the amount of items detected per platform increase as you browse. Toggle
collection off when done to avoid inadvertantly gathering undesired data.

When you have the items you need, you can export the data as an [ndjson](https://ndjson.org) file, or upload
it to a 4CAT instance where a 4CAT dataset will be created from the uploaded items. You can then run 4CAT's analytical
processors on the data.

To upload to 4CAT, copy the URL of the website of the 4CAT instance to the "4CAT instance" field at the top of
Zeeschuimer's interface. You can then use the "to 4CAT" button to create a new 4CAT dataset from the captured data.
**To upload to 4CAT, copy the URL of the website of the 4CAT instance to the "4CAT instance" field at the bottom of
Zeeschuimer's interface**. You can then use the "to 4CAT" button to create a new 4CAT dataset from the captured data.
After uploading, Zeeschuimer will show you a link and the ten most recently uploaded datasets are shown at the bottom of
the interface.

Expand All @@ -77,6 +90,13 @@ platform. The following limitations are known:
* *TikTok* items that cannot be captured:
* Live streams

For some platforms, the level of detail of the data that can be collected depends on the page it is captured from:

* *Pinterest* items may lack some metadata unless captured from the individual post's page, most notably the timestamp
of the post.
* *RedNote/Xiaohongshu* items will often lack the item's post description, timestamp, and video URL, unless captured by
opening the post's own page/clicking it in an overview.

Note that these are *known* limitations; data capture may break or change based on platform changes. Always
cross-reference captured data with what you are seeing in your browser.

Expand Down
Binary file added images/platform-icons/pinterest.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added images/platform-icons/xiaohongshu.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added images/zeeschuimer-16.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
23 changes: 23 additions & 0 deletions js/lib.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
function traverse_data(obj, callback) {
let results = [];

function _traverse_data(obj, callback) {
for (const property in obj) {
if (!obj.hasOwnProperty(property) || !obj[property]) {
// not actually a property
continue;
}

let callback_result = callback(obj[property], property);

if (callback_result) {
results.push(callback_result);
} else if (typeof (obj[property]) === "object") {
_traverse_data(obj[property], callback);
}
}
}

_traverse_data(obj, callback);
return results;
}
13 changes: 10 additions & 3 deletions js/zs-background.js
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ window.zeeschuimer = {

// the document can be parsed by all modules listening on either the origin or document's URL's domain
let eligible_modules = Object.fromEntries(Object.entries(window.zeeschuimer.modules).filter(entry => {
return possible_source_domains.includes(entry[1]["domain"].toLowerCase());
return possible_source_domains.some((domain) => domain.endsWith(entry[1]["domain"].toLowerCase()));
}));

filter.ondata = event => {
Expand All @@ -95,15 +95,17 @@ window.zeeschuimer = {

filter.onstop = async (event) => {
// pass the document to all eligible modules that are also enabled
let enabled_modules = [];
for(const module_id in eligible_modules) {
const module_enabled_key = 'zs-enabled-' + module_id;
let module_enabled = await browser.storage.local.get(module_enabled_key);
module_enabled = module_enabled.hasOwnProperty(module_enabled_key) && !!parseInt(module_enabled[module_enabled_key]);

if(module_enabled) {
await zeeschuimer.parse_request(full_response, origin_url, document_url, details.tabId);
enabled_modules.push(module_id);
}
}
await zeeschuimer.parse_request(full_response, origin_url, document_url, details.tabId, enabled_modules);
filter.disconnect();
full_response = '';
}
Expand All @@ -117,8 +119,9 @@ window.zeeschuimer = {
* @param origin_url URL of the *page* the data was requested from
* @param document_url URL of the content that was captured
* @param tabId ID of the tab in which the request was captured
* @param enabled_modules List of IDs of enabled modules
*/
parse_request: async function (response, origin_url, document_url, tabId) {
parse_request: async function (response, origin_url, document_url, tabId, enabled_modules) {
if (!origin_url) {
origin_url = document_url;
}
Expand Down Expand Up @@ -160,6 +163,10 @@ window.zeeschuimer = {

let item_list = [];
for (let module_id in this.modules) {
if(!enabled_modules.includes(module_id)) {
continue
}

item_list = this.modules[module_id].callback(response, origin_url, document_url);
if (item_list && item_list.length > 0) {
await Promise.all(item_list.map(async (item) => {
Expand Down
15 changes: 12 additions & 3 deletions manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -3,12 +3,15 @@
"description": "Collect data while browsing social media platforms and upload it for analysis later",
"manifest_version": 2,
"name": "Zeeschuimer",
"version": "1.11.1",
"version": "1.13.4",
"homepage_url": "https://github.com/digitalmethodsinitiative/zeeschuimer",

"browser_specific_settings": {
"gecko": {
"update_url": "https://extensions.digitalmethods.net/updates.json"
"update_url": "https://extensions.digitalmethods.net/updates.json",
"data_collection_permissions": {
"required": ["none"]
}
}
},

Expand All @@ -33,6 +36,7 @@
"scripts": [
"inc/dexie.js",
"inc/he.js",
"js/lib.js",
"js/zs-background.js",
"modules/tiktok.js",
"modules/tiktok-comments.js",
Expand All @@ -44,7 +48,12 @@
"modules/douyin.js",
"modules/gab.js",
"modules/truth.js",
"modules/threads.js"
"modules/threads.js",
"modules/pinterest.js",
"modules/rednote.js",
"modules/rednote-comments.js",
"modules/facebook.js",
"modules/facebook-comments.js"
]
}
}
Loading