Thread
Original post has been removed.
The models always used internal reasoning; predicting the next token was simply the output method. Some people experimented to see if the models could forecast 2 or 3 tokens ahead in the output, and they can.
ago
0 Likes0 Dislikes0 Replies
?
No replies yet