Introducing TriTopic: How Multi-View Graphs and Consensus Clustering solve the biggest flaws of modern NLP models.
The automatic structuring of text data is the backbone of almost every modern Natural Language Processing (NLP) pipeline. When faced with hundreds of thousands of customer reviews, complex research abstracts, massive legal corpora, or daily news articles, human categorization is no longer feasible.
The 3 Great Illusions of Current Topic Models
1. The Outlier Illusion (The 20% Data Loss)
Density-based clustering algorithms like HDBSCAN discard documents as "noise." BERTopic discards an average of nearly 20% of documents as outliers. In legal discovery, medical records, or intelligence work, ignoring a fifth of your data is unacceptable.
2. The Stability Dilemma
Stochastic algorithms are inherently unstable. Change the random seed and the entire topic structure shifts. For rigorous scientific studies and reliable BI systems, a model lacking reproducibility is useless.
3. The "Embedding Blur"
Sentence-BERT compresses semantically rich documents into a single vector, losing lexical precision in the process. The result: topics that are vague and hard to interpret.
Introducing TriTopic
TriTopic addresses all three weaknesses through a tri-modal graph that fuses semantic embeddings, TF-IDF weights, and metadata. Three core innovations drive its performance:
- Hybrid graph construction via Mutual kNN and Shared Nearest Neighbors to eliminate noise
- Consensus Leiden Clustering for reproducible, stable partitions
- Iterative Refinement that sharpens embeddings through dynamic centroid-pulling
In benchmarks across 20 Newsgroups, BBC News, AG News, and Arxiv datasets, TriTopic achieves the highest NMI on every dataset (mean NMI 0.575 vs. 0.513 for BERTopic) and guarantees 100% corpus coverage with 0% outliers.
TriTopic is available as an open-source PyPI library. Read the full paper on ArXiv →
By Prof. DDr. Roman Egger