Add generation markers to the chat template
This PR adds {% generation %} / {% endgeneration %} markers around the assistant output in chat_template.jinja, so return_assistant_tokens_mask=True returns the assistant tokens. Without them the mask is all zeros, and assistant-only loss (e.g. TRL SFT with assistant_only_loss=True) can't work. Rendering is unchanged.
The <|im_start|>assistant\n header is now emitted once before the branches, so it stays outside the mask. The mask covers what the model generates, up to and including <|im_end|>:
<think>\nr\n</think>\n\nHello<|im_end|>\n
Checked: the rendered prompt is identical before and after for 40 combinations (system, multi-turn, tool calls and responses, with and without tools, add_generation_prompt, enable_thinking). The mask covers each assistant turn in multi-turn and tool-call conversations, and the <|im_end|> stop token is always inside it.