Usual model/quantization analysis involves computing the perplexity of the raw LLM logits on reference text. These tools are for a slightly different purpose: looking at the perplexity under sampling conditions, i.e., after the sampling pipeline has been applied. The goal is to allow for a more informed choice of sampling parameters and to more easily test the impact of new sampling methods.
The basic mechanism of sampling in large language model (LLM) is as follows:
- logits = LLM(tokenized preceding text)
- logits = filter(logits)
- probs = softmax(logits/Temperature)
- token ~_{sample} probs
The resulting token is then appended to the previously generated tokens, and the steps repeat.
In the filtering step, tokens are discarded following various rules:
- top_k discards all but the top k most likely tokens;
- grammars discard all tokens that do not match the required grammar;
- min_p discards all tokens which, depending on the min_p threshold, are too unlikely compared to the most likely token.
In the final sampling step, a token is selected randomly with probability proportional to the computed probability of the token.
The basic log-likelihood is $$ ll = \sum_i logits[i][correct_token[i]] $$, where the logits[i] are the computed logits for a text at token i, and correct_token[i] is the actual correct token of the reference text. The perplexity is simply the exponential exp(-ll).
The sampling log-likelihood then is, naively, $$ ll_s = \sum_i \log(probs[i][correct_token[i]]). $$ The issue is that the probability of a correct token can (and often will) be 0, leading to -infinity. The solution is to clip: $$ ll_s = \sum_i \log(\max(probs[i][correct_token[i]], \epsilon)). $$ The parameter epsilon>0 simply treats any text token with a probability smaller than epsilon as having probability epsilon. This makes sense for the sampling log-likelihood, since on the sampling side of things, it does not matter if the correct token had probability e-4 or e-5, the chance that during sampling the correct token is chosen is so small as to be negligible (this is different to a training scenario, where that difference might be a relevant training signal). For this reason, I chose in all the analysis epsilon = 0.01.
The important thing is that ll_s is a valid optimization target, allowing us to compare the influence of the sampling pipeline on ll_s for a single model/quantization.
The optimization goal then is to minimize the negative sampling log-likelihood. A side target might be to minimize the number of out-of-distribution tokens, which here is the number of tokens whose probability was replaced by epsilon.
In the sampling pipeline, tokens are sampled proportionally to the computed probs. This is the same as subdividing the unit interval [0,1] into subintervals of length probs[i], choosing a point on [0,1] uniformly at random, and checking in which interval the chosen point landed: this represents the selected token. In the sampling pipeline the candidate tokens where already ordered according to likelihood, so in this representation the largest subinterval corresponding to the most likely token is the left-most interval, the second likeliest token is the second subinterval, and so on, with all the unlikely tokens clustered at the right of [0,1]. Since the goal of choosing temperature and cutoff strategies like min_p sampling to gain control over how to best over-/underrepresent tokens with large, medium or small log-likelihood there is a natural way to generalize the sampling step: Instead of sampling a point uniformly at random on [0,1] to determine the selected token, we can sample the point from a beta-distribution.
The beta-distribution is a two-parameter family of distributions on [0,1] with parameters (alpha, beta) which generalizes the uniform distribution. For alpha=beta=1 it is in fact the uniform distribution. The parameters alpha and beta determine the behaviour of the distribution near 0 and 1 respectively. For 0<alpha<1, it is more likely to sample near 0, while for alpha>1 it is less likely. The same is true for beta and sampling near 1.
So, if one wants to increase the likelihood of the most likely token, one could decrease the temperature, or one could decrease the alpha value. Similarly, if one would like to increase the chance to sample moderately likely tokens, increasing both alpha and beta would penalize both the most likely token and the low-probability tokens.
scripts/download_datasets_cli.py will download a huggingface dataset or read a parquet file, and turn it into a set of prefix/continuation .txt files in the targeted folder. prefix files are used to provide context to the model, while continuation is used to compare LLM predictions with the actual text.
scripts/LogProbs.py will use a given dataset and a llamacpp server endpoint to compute the token logprobs and alternatives. Uses the correct output text (continuation.txt) as a grammar string literal to obtain token logits for the entire text in a single API call. Will default to the top 100 token alternatives per text token. Saves result in .npz file. Can loop over all models the server provides if --all_models is set.
scripts/analyze_logprobs_cli.py will read a .npz file and try to find the optimal temperature and/or the optimal (alpha,beta) for a beta-distribution by minimizing the negative sampling log-likelihood. By default it will try multiple random starts to improve the chance of finding the global minimum. In experiments usually all starts converge to the same optimium, but sometimes this fails. Potential reasons are:
- ll_s is piecewise constant as min_p changes; gradient methods are problematic, optimization over min_p disabled by default;
- with min_p>0, there is a competition between two objectives: minimizing the in-distribution negative log-likelihood, and minimizing the out of distribution percentage;
- with min_p>0, the negative sampling log-likelihood is not convex. Results can be written into a .csv for further analysis.
scripts/batch_logprobs.ps1 is a wrapper around scripts/analyze_logprobs_cli.py to do multiple analysis runs over multiple .npz and settings
scripts/plot_logprobs.py will read a .csv generated by the above scripts and plot the negative sampling log-likelihood and out of distribution percentage, grouped by model/dataset.
scripts/param_sweep.py will read a .npz file and do a simple temperature sweep from 0 to 2.5 and plot the results.
scripts/batch_param_sweep.ps1 will read all .npz files in a directory and run param_sweep.py over all those files and, optionally, various parameter settings