This repository implements a pipeline to store various data of files from a large unstructured dataset. These fields are used for topic modeling (wordclouds, based on low-dimensional versions of embedding vectors, Named Entity Clustering and document-topic incidences). The information is aggregated and visualised using FCA.
elasticsearch visualisation embeddings documents ner text-data fca topics-modeling sentence-transformers top2vec topic-aggregation ner-clustering
-
Updated
Jan 17, 2026 - Python