> Is it okay to train AI on your assets and/or modify your assets using AI?
I don't think this is something that OGA has any jurisdiction over; only the licenses do, and the licenses grant permission for any use, even disagreeable ones (which is a core tenet of the libre philosophy since the 20th century, I'm afraid to say). Sketchfab is in a similar situation, and here's what they had to say:
> If a creator applies the NoAI tag to a model they’ve made available under a Creative Commons license, the tag may not be enforceable outside of the Sketchfab platform. This means that for CC-licensed work, the NoAI tag might not be respected because the CC license may allow use of the work by generative AI programs regardless of the tag. This is due to the way CC licenses work and is not something Sketchfab can control. https://sketchfab.com/blogs/community/restricting-generative-ai-use-of-free-models/
Creative Commons themselves are unlikely to add a restriction in, say, version 5.0 of their licenses, given their official viewpoint is that training generative models on even wholly unlicensed data is fair use, and their recommendation for those who don't want that to happen is Asking Nicely.
I will note that, if it's decided that models are non-fair·use derivative works of their training data, the attribution requirement in almost all licenses (except CC0) will mean it's still an infringement, unless you catalog and credit every piece of data that went in (not impossible!). CC0 would still be trainable on, unless the governments of the world decided to establish an entirely new opt-in right for AI training (a la moral rights), which at least in the current US climate feels unlikely (though there have been stranger carve-outs).
---
> while still claiming his game uses real artists' art for PR purpose
Hmm, "contains human assets" vs "doesn't contain generated assets"... For whatever it's worth, both Steam and itch.io expect you to disclose the existence of any of the latter. Either way, I'd imagine style mimicry would still be looked down upon and called out, even in your scenario.
I don't think the FSF will ever approve of a license that excludes training, since it goes against the zeroth freedom: "the freedom to use the program {or assets, in this case} for any purpose". Heck, the official opinion of Creative Commons is that training generative models on copyrighted material is likely to be fair use.
Either way, most currently available licenses require credit, which no mainstream image generator gives, so the point's kind of moot. The best one could do, in my opinion, is put up a tiny barrier that would discourage most scrapers, e.g. uploading as a .zip file (perhaps with a simple password). Unless the concern is individuals copying the style, which... hmm, I'm not sure if any sequence of words would discourage them.
Looking forward to when/if there is a fully openly-trained model some day.
There's one in the works called Public Diffusion. Much of the data is taken from Wikimedia Commons, so it does have some problematic images dotted about, whether due to quirks of the site (cosplay of copyrighted characters is allowed for some reason), differences in copyright terms across countries, or just blatant copyright infringement that didn't get caught. They also use a scrape-trained language model for captioning and another for interpreting the captions, which may or may not matter copyright-wise.
(Another project, Elan Mitsua, is stricter on both counts, but the terms of use are likely too strict for OGA (and aside from public-domain images, it's also trained on works submitted specifically for training). It's also not quite there in terms of quality for those who want to generate ready-to-use assets, but it can serve as inspiration, if nothing else.
> Is it okay to train AI on your assets and/or modify your assets using AI?
I don't think this is something that OGA has any jurisdiction over; only the licenses do, and the licenses grant permission for any use, even disagreeable ones (which is a core tenet of the libre philosophy since the 20th century, I'm afraid to say). Sketchfab is in a similar situation, and here's what they had to say:
> If a creator applies the NoAI tag to a model they’ve made available under a Creative Commons license, the tag may not be enforceable outside of the Sketchfab platform. This means that for CC-licensed work, the NoAI tag might not be respected because the CC license may allow use of the work by generative AI programs regardless of the tag. This is due to the way CC licenses work and is not something Sketchfab can control.
https://sketchfab.com/blogs/community/restricting-generative-ai-use-of-free-models/
Creative Commons themselves are unlikely to add a restriction in, say, version 5.0 of their licenses, given their official viewpoint is that training generative models on even wholly unlicensed data is fair use, and their recommendation for those who don't want that to happen is Asking Nicely.
I will note that, if it's decided that models are non-fair·use derivative works of their training data, the attribution requirement in almost all licenses (except CC0) will mean it's still an infringement, unless you catalog and credit every piece of data that went in (not impossible!). CC0 would still be trainable on, unless the governments of the world decided to establish an entirely new opt-in right for AI training (a la moral rights), which at least in the current US climate feels unlikely (though there have been stranger carve-outs).
---
> while still claiming his game uses real artists' art for PR purpose
Hmm, "contains human assets" vs "doesn't contain generated assets"... For whatever it's worth, both Steam and itch.io expect you to disclose the existence of any of the latter. Either way, I'd imagine style mimicry would still be looked down upon and called out, even in your scenario.
I don't think the FSF will ever approve of a license that excludes training, since it goes against the zeroth freedom: "the freedom to use the program {or assets, in this case} for any purpose". Heck, the official opinion of Creative Commons is that training generative models on copyrighted material is likely to be fair use.
Either way, most currently available licenses require credit, which no mainstream image generator gives, so the point's kind of moot. The best one could do, in my opinion, is put up a tiny barrier that would discourage most scrapers, e.g. uploading as a .zip file (perhaps with a simple password). Unless the concern is individuals copying the style, which... hmm, I'm not sure if any sequence of words would discourage them.
There's one in the works called Public Diffusion. Much of the data is taken from Wikimedia Commons, so it does have some problematic images dotted about, whether due to quirks of the site (cosplay of copyrighted characters is allowed for some reason), differences in copyright terms across countries, or just blatant copyright infringement that didn't get caught. They also use a scrape-trained language model for captioning and another for interpreting the captions, which may or may not matter copyright-wise.
(Another project, Elan Mitsua, is stricter on both counts, but the terms of use are likely too strict for OGA (and aside from public-domain images, it's also trained on works submitted specifically for training). It's also not quite there in terms of quality for those who want to generate ready-to-use assets, but it can serve as inspiration, if nothing else.
here's a text-to-speech engine with a bunch of voices under various cc licenses, some foss-compatible: https://github.com/rhasspy/piper
(though most of them were trained by starting from the model for "lessac", which has a restrictive research license; idk how much that matters though)
and here's an asset pack made with it: https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to-speech