java statistics lda topic-modeling mallet

Strange perplexity values of LDA model trained with MALLET

I have trained an LDA model with MALLET on parts of the Stack Overflow data dump and did a 70/30 split for training and test data.

But the perplexity values are strange, because they are lower for the test set than for the training set. How is this possible? I thought the model is better fitted for the training data?

I have already double checked my perplexity calculations, but I do not find an error. Do you have any idea what the reason could be?

Thank you in advance!

Edit:

Instead of using the console output for the LL/token values of the training set, I have used the evaluator on the training set again. Now the values seem to be plausible.

Solution

That makes sense. The LL/token number is giving you the probability of both topic assignments and the observed words, whereas the held-out probability is giving you the marginal probability of just the observed words, summed over topics.

Keep numbers which appear in both columns, in J lang
Differentiation in J
How to reshape an array with an arbitrary size in one dimension?
Why is Insert (fold) right associative
Write 4 : 'x&{.&.;: y' tacitly
Alignment issue when printing formatted prime numbers in J language
How can I define a verb in J that applies a different verb alternately to each atom in a list?
How to get user input in the J programming language
How to unbox a list of boxed lists of differing lengths in J?
How can I fix 'noun result was required' error in J?
In j, how can I define a verb locally in one scope and pass it to a defined adverb?
Convert boxed array to normal array?
Read column of CSV file as array
Replace atom in array of strings
What does the dyad `=` do to boxed strings?
Index of minimum element using J
How can I take the outer product of string vectors in J?
Building an array of verbs in J
Reading in multidigit command line parameter
Amend with bond to new data shows unexpected behaviour
How to turn a table or matrix into a (flat) list in J
How to run dissect in J?
How to define selection using index function in J
How to exit the J console?
Find 4-neighbors using J
Writing custom verbs in J
How do I negate a selector in J lang?
How to use arbitrary selector in interchange in J lang?
different result once square root is added inside tacit
Sum of arrays with repeated indices