The sentience argument is a bit confusing. The fact that it can produce language is definitely interesting but with that said I haven't seen any arguments that stable diffusion is sentient.
The technology although different is also mostly the same.
Can someone provide me why image generation is not sentient but word generation is?
Because humans are much more strongly linguistic than artistic.
But also because so far, LLMs don't generate images in reaction to prompts. They generate images that match prompts. But that is probably easy to fix by a clever ML person.
Imagine a perfect simulation of a brain within a computer. It can learn/feel/remember/act/perdict/etc. just like a human brain. Most people would call this simulation sentient. If you took that same simulation but changed the internal structure to something else, but it was still capable of the exact same outputs for given inputs, is it still sentient? A perfect chatbot would be an example of one of these simulations.
Your answer to the question isn't really the point here, i'm just trying to clarify the argument.
The technology although different is also mostly the same.
Can someone provide me why image generation is not sentient but word generation is?