‹ Watch
Method

How your chills score is computed

Your score is a rank. You answered three short questionnaires. Those answers go into a formula that was fitted on 2,937 people who answered the same questions and watched the same videos in a peer-reviewed study. The formula returns your position on their distribution. Every step is on this page.

1What we measure

Before you watched, you answered a short set of questions. They come from three instruments that decades of research connect to aesthetic chills — the shivers, goosebumps, and waves of cold that some experiences trigger:

Positive emotionHow strongly everyday moments of awe, beauty, and joy land on you. Three items.Dispositional Positive Emotion Scales (DPES)
PersonalityFive short items from the standard five-factor personality inventory. Five items.NEO Five-Factor Inventory (NEO-FFI)
Being movedHow easily you're touched — by reunions, kindness, music, farewells. One item.Kama Muta Frequency scale (KAMF)

2The reference sample

In 2023 our team collected the same measures from 3,259 people in Southern California; 2,937 remained after quality checks. Each watched videos drawn from a validated set of 40 chills-eliciting stimuli, then reported whether chills happened and how intense they were. The dataset — ChillsDB 2.0 — is published open access, so any researcher can rerun what we do here.

3How the model learned

This is where machine learning comes in, and the idea underneath it is simple arithmetic. Every participant becomes one row of numbers. Eleven of those numbers are features — the things known before the video played: nine questionnaire answers, age, and sex. One more thing the model sees is which video was playing, because some videos give chills to almost everyone and some to almost no one. And one number is the label — what actually happened: chills or no chills.

The model is a set of weights, one per feature, each saying how much that feature counts. Training runs a loop: read a row, multiply each feature by its current weight, add the results, convert the total to a probability, compare that probability to the real label, and nudge every weight a little in whichever direction would have made the guess closer. Repeat across all 2,937 rows, many times over. The weights drift until the error stops shrinking, and that resting point is the model — a logistic regression, the most transparent model in the machine learning toolbox: every prediction is a weighted sum you could check by hand. Here is roughly where the weights landed:

KAMF
strongest
DPES
strong
NEO-FFI
moderate
Age
small
Sex
small

Relative weight each block of features carries in the fitted model. How easily you are moved dominates; age and sex fine-tune.

4Why we test on strangers

A model with enough freedom can memorise the very people it was fitted on and then fail on anyone new. Researchers call that overfitting, and the standard guard against it is cross-validation. The 2,937 participants get split into ten groups. The model trains on nine of them and is scored on the tenth, which it has never seen. Rotate ten times, so every participant serves once as a stranger, and average the ten scores.

Round 1
Round 2
Round 3
…
trained on scored on

Every accuracy figure on this page comes from those held-out rounds.

5From your answers to a percentile

Your answers→ Scored on each scale→ Weighted by the model→ Ranked against 2,937→ Your percentile

Each questionnaire is scored by its own published rules, the eleven numbers go through the weights above for every one of the 40 videos, and your average predicted chills probability lands somewhere on the curve built from the reference sample. A score of 88% means you're more likely to get chills than 88% of those participants. "1 in 8" counts the share at your level or above — the same fact, said the other way.

The chart on your profile is this distribution — every bar is a slice of the 2,937, and the marked bar is where you land.

6How well it performs

On held-out participants the model reaches an AUC of 0.767. In plain terms: hand it one person who got chills from a video and one who did not, and it ranks the two correctly about 77 times out of 100. A coin flip would manage 50. The probabilities are also calibrated — when the model says 60%, chills happen roughly 60% of the time — which is what lets us turn them into an honest score. It carries real signal, and it also leaves plenty on the table.

A higher score means chills are more likely for you, on average, across the kinds of videos in the study. It describes a propensity, not a promise — the right video on the right day matters as much as the trait.

7The limits

Read the score as one narrow statement: how your questionnaire answers rank against this particular sample. Its edges are worth knowing. The reference group comes from one region, Southern California, at one moment in time. Chills reports are self-reported. And an AUC of 0.767 means the ranking goes the wrong way roughly one time in four. We show you the arithmetic because the honest version is interesting enough.

8Sources

Predicting individual differences in peak emotional response. Schoeller, Christov-Moore, Lynch, Diot & Reggente. PNAS Nexus, 2024. doi.org/10.1093/pnasnexus/pgae066
ChillsDB 2.0 — the open dataset (CC BY 4.0). Schoeller et al., Scientific Data, 2023. Download on FigShare
ChillsDB: a gold standard for aesthetic chills stimuli. Schoeller et al., Scientific Data, 2023.

ChillsTV is for education and research purposes only. It provides no medical or psychological advice. Questions: [email protected]

Copied