I don’t think there’s any legal issue with using data for training. I think the problem is how do you subsequently prevent your model from creating copyright violations. If you watch a Spider-Man movie and take some inspiration from it, you know that you go and make a movie with similar themes or whatever, but you can’t just go out and create your own Spider-Man movie. An AI model doesn’t know this, and I don’t know how you’d teach it this concept. Especially when properly educated humans frequently have disputes about what is/isn’t allowed.
Spiderman was released 1962. Copyright runs out in 2057.
What you could do is what some of of the original comic book creators did when they ripped off each others work (Doom Patrol/Xmen, Quicksilver/Flash, etc.) - create an alternative version that was close the original but was different enough that it wasn't a straight up copy. So it wouldn't be Spiderman, but Venom or something, and the suit wouldn't be red/blue but like maybe purple/green or something.
Most likely these models are going be hidden behind the network. Aside from the issues of size, which will probably run into the terabytes, I can't see companies willing to risk including them in downloadable code them given how expensive they are to generate, for fear of being copied.
A single network transfer alone is enough to qualify "sharing".
Heck some of the content creators are even arguing that the models are essentially a form of lossy compression, because no one quite knows exactly how they work.
Also the users probably won't be the ones with deep pockets, unless it was a studio. And the lawyers are generally going to after the ones with deep pockets.