The provider looks up previously computed key/value tensors. Prefill starts after the match. Those tokens are billed as cached input — cheaper, and time-to-first-token drops, especially on long prompts.
The generated text is still new. A prompt that “means the same thing” but differs by one character in the prefix is a miss. Semantic caches in your app are a different layer entirely.