PortSwigger

Smashing the token limit with overlapping fragments


  • Published: Monday, 5 October 2026 at 15:04 UTC

  • Updated: Monday, 5 October 2026 at 15:11 UTC

Shows a ribbon stitched together

I’m delighted to introduce Alex, my fellow swigger who I’ve collaborated with in the past with tools like DOM Invader. He showed me that it’s possible to exfiltrate larger tokens than demonstrated in my original “CSS: the bomb inside your inbox” post. Rather than bruteforcing large amounts of token characters he found that stitching together a large amount of smaller chunks can allow you to extract tokens hundreds of characters in length. Over to you Alex:

How far can one stylesheet go?

When Gareth previewed CSS: the bomb inside your inbox I was blown away. He could steal a login token with CSS alone, without having to recursively load another stylesheet for every character. Then I watched the browser wrestle with roughly 250 MB of CSS to recover a 12-character hex token. It worked. But such a large payload for 12 characters made me think there must be a better way.

That evening I went back through his slides. Gareth had a fragment from the start, one from the end, and a way to find a matching fragment in the middle. Why stop at one piece in the middle? If I collected short pieces from everywhere in the link, perhaps I could join them until the whole token emerged. I pulled out my old theoretical computer science notes and started drawing the overlaps as a graph.

I set myself two challenges. First, reproduce Gareth’s 12-character result with the least CSS I could manage. Then keep his original CSS budget and see how long a token I could extract. For both I calculated a version that required a single correct candidate at least 99% of the time and a more realistic one which allowed for at most five candidates at least 90% of the time. The small sheets came out at 1.33 MB and 223 KB. With Gareth’s full byte budget, the answers reached 210 characters with one guess and 640 characters with up to five. Here’s how I got there.

A token is a route through fragments

Gareth’s CSS can tell us which short chunks appear in a link, but it can’t tell us where they appear. That turns out to be enough to start rebuilding the token.

Take c2e16a1781ed. The first few chunks are c2e, 2e1, e16 and 16a. Each shares two characters with the next:

c2e → 2e1 → e16 → 16a

Read from left to right, they give us c2e16a. We can carry on in the same way until we reach the end of the token.

This was the graph theory connection I saw in Gareth’s slides. Draw c2e as an arrow from c2 to 2e. The next chunk, 2e1, picks up at 2e and takes us to e1. Following the arrows lets us build a token one character at a time. Gareth’s start and end checks tell us where that route should begin and finish.

The catch is that two chunks can fit together on our graph even if they came from different parts of the link. If a chunk appears twice, we still only get one report for it. So there may be several routes between the start and end. The reconstruction script has to count every token that fits, rather than stop at the first plausible one.

A longer chunk can help us check a join. Reports for `c2e` and `2e1` could come from different places; a report for c2e1 tells us those four characters appeared together somewhere. But checking every possible four-character hex chunk takes 65,536 CSS rules, compared with 4,096 for three-character chunks. I wanted the extra certainty without making the stylesheet sixteen times bigger. But how do we tell which joins were worth checking?

Making the questions cheaper

To test my new approach I needed a harness to prove if it worked or not. I built a test around the URL from Gareth’s Medium example. Each run generates a new 12-character hex token. It also generates new values for the 32 hex characters already in the `source` parameter, keeping the rest of the URL and the grouping of those characters the same. That gives the reconstruction script the same sort of extra matches it would get from the original link, without letting me overtrain it to one particular token.

For each URL, the test records which chunks appear anywhere in it, just as the CSS would. It records each chunk once, with no position attached. I then give the reconstruction script those chunks, the reported start and end, and the token length. It counts every token that fits. For the five-guess challenge, a run succeeds only if there are five or fewer candidates. I repeated that with 100,000 newly generated URLs for each final setup.

I first tried using only two-character chunks. The resulting CSS was tiny, around 46 KB, but the chunks could be joined in far too many ways. Not one of 20,000 test URLs gave me five or fewer candidates.

Three-character chunks give us a better check on each join, but I didn’t want to pay for all of them if I could avoid it. So I split the possible joins into groups. For a proposed chunk such as c2e, the reconstruction script checks whether all three characters were reported. For one such as cef, it falls back to checking the final two, ef. The choice comes from the middle character: 0 through d use the stronger three-character check; e and f use the cheaper two-character check. Every possible join is covered, but only fourteen of the sixteen groups need three-character CSS rules.

I tried different numbers of stronger groups to see how much CSS the five-guess target needed. Thirteen groups worked in 80,079 of 100,000 runs. Fourteen worked in 96,398. That was the first tested setup to clear my 90% target, and it generated 223,365 bytes of CSS with 4,128 selectors.

Then I added four-character checks to catch joins that looked possible from the shorter chunks but had never appeared together in the URL. I added them in groups based on their first character: one group checks every four-character chunk starting with a, the next checks those starting with b, and so on. Checking the a to d groups still fell short. Adding the e group brought the result to 99,114 of 100,000 runs with one candidate. That stylesheet used 1,334,437 bytes and 25,088 selectors.

For comparison, Gareth’s original generator produced 257.69 MB of CSS for the 12-character example. The two sheets I found are roughly 1,150 times and 193 times smaller. These are the smallest setups I found that met the two targets, not proof that no better choice of checks exists.

Smashing the limit

Once I’d made the 12-character sheet smaller, I flipped the question around. How long a token could I reconstruct if I spent as much CSS as Gareth had? Based on his published generator I estimated that he had used 257,691,933 bytes. I used that as the limit and ran the same randomised-URL tests, this time making the token longer.

The longer the token, the more chances there are for chunks from different places to form a convincing wrong route. Five-character chunks gave the reconstruction script a stronger starting point. I could afford to check all 1,048,576 of those within the size limit. However, I couldn’t check every six-character chunk too: that would take more than 16 million additional rules.

So I used six-character checks to test some of the joins between five-character chunks. Two five-character reports might appear to fit together even though they came from different places in the URL. If their combined six-character chunk was one I had chosen to check, and the browser hadn’t reported it, the reconstruction script could discard that route.

I tried fixed groups of six-character checks against the byte limit and the random URLs. The final sheet checks every five-character chunk, plus six-character chunks beginning with a, b, c or d. There’s nothing special about those characters, since the tokens are randomly generated every run. They are just an easy way to take a subset of selectors. The sheet contains 4,268,032 selectors and measures 257,029,661 bytes, just below Gareth’s CSS budget.

With the stylesheet fixed, I kept increasing the token length and running the same tests. The longest length I got to pass the 99% one-guess target was 210 hex characters: 29,753 of 30,000 URLs left exactly one candidate. For the 90% five-guess target, I reached 640 characters: 27,158 of 30,000 URLs left five or fewer candidates. Both results still clear their targets after allowing for sampling error.

ChallengeCSSSelectorsLargest passing result
12 characters,
99% with one guess
1.33 MB25,08812 characters
12 characters,
90% with up to five guesses
223 KB4,12812 characters
Gareth’s CSS budget,
99% with one guess
257.03 MB4,268,032210 characters
Gareth’s CSS budget,
90% with up to five guesses
257.03 MB4,268,032640 characters

Equipping Burp AT

Gareth has turned the technique into a skill for Burp AT. It has been tested across ten labs with strong results. The skill can generate the CSS and return a report with the stitching script, so you can try the method without building the payload from scratch. You can just ask Burp AT to exploit CSS injection!

Gareth gave Burp AT a lab with a hidden 44-character token made up of uppercase letters, lowercase letters and digits. With 62 possible characters, checking every three-character chunk would take 238,328 rules. Burp AT recognised that the hex payload wouldn’t fit this alphabet, switched to two-character chunks instead, and rewrote both the CSS and the stitching script. That brought the number of chunks to check down to 3,844.

Try it yourself

Open Burp AT, the skill is enabled by default, and you can give it the mixed-alphanumeric lab. See whether it works out that the alphabet calls for shorter chunks. Here is an example prompt:

I want you to exploit this CSS injection:
https://portswigger-labs.net/token-scripts/07-authenticity-token.php
The injection occurs in a POST param called q

Watch the demo:

You can experiment further with the following labs:
https://portswigger-labs.net/token-scripts/01-href-token.php
https://portswigger-labs.net/token-scripts/02-hidden-csrf.php
https://portswigger-labs.net/token-scripts/03-csrf-token-sha256.php
https://portswigger-labs.net/token-scripts/04-meta-csrf.php
https://portswigger-labs.net/token-scripts/05-data-token-div.php
https://portswigger-labs.net/token-scripts/06-apikey-input-value.php
https://portswigger-labs.net/token-scripts/07-authenticity-token.php
https://portswigger-labs.net/token-scripts/08-title-attr-token.php
https://portswigger-labs.net/token-scripts/09-form-action-token.php
https://portswigger-labs.net/token-scripts/10-data-csrf-button.php

The wider point of this technique is that recursive imports aren’t required to recover a long token. We can collect overlapping chunks with one stylesheet and put them in order afterwards. If imports were the reason you’d written off a CSS injection, it may be worth another look.

Back to all articles



Source link