Synthetic realities

Generative video: creativity, epistemic limits and the verifiability of communication

Keywords: AI Act, Foundation models, Generative video, Latent diffusion, Deepfake

Abstract

This paper offers a critical, interdisciplinary review of the current state of the art in generative artificial intelligence (AI), focusing on the transition from linguistic models to generative video models and its implications for public communication. The central argument is that the scaling paradigm — the idea that an increase in parameters, data and computing power produces predictable improvements — has made it possible, in less than four years, to synthesise photorealistic audiovisual sequences, but is now constrained by three types of limitations: (i) epistemic limits, relating to the measurement of capabilities, the physical consistency of generated scenes and the statistically inevitable nature of hallucinations; (ii) communicative limits, relating to the erosion of the probative value of audiovisual recordings; (iii) normative limits, relating to the emergence of binding AI law in Europe. It is shown how the ‘world simulator’ thesis (OpenAI 2024) is refuted by physical common-sense benchmarks (Bansal et al. 2024; Kang et al. 2024), such as the liar’s dividend (Chesney and Citron 2019) and the loss of the ‘epistemic support’ provided by the recording (Rini 2020) redefine the problem of disinformation, and how Article 50 of Regulation (EU) 2024/1689 and Italian Law 132/2025 translate these technical issues into legal obligations. The analysis concludes with a research agenda focused on evaluation, the provenance of content, trust and the impacts on the creative industries.

References

- Acemoglu, D. (2024). The simple macroeconomics of AI (NBER Working Paper No. 32487). National Bureau of Economic Research. https://doi.org/10.3386/w32487
- Alibaba (Wan-Video Team). (2025). Wan: Open and advanced large-scale video generative models [Technical report].
- Bansal, H., Lin, Z., Xie, T., Zong, Z., Yarom, M., Bitton, Y., Jiang, C., Sun, Y., Chang, K.-W., & Grover, A. (2024). VideoPhy: Evaluating physical commonsense for video generation. arXiv. https://arxiv.org/abs/2406.03520
- Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922
- Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., & Rombach, R. (2023). Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv. https://arxiv.org/abs/2311.15127
- Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
- Chesney, R., & Citron, D. K. (2019). Deep fakes: A looming challenge for privacy, democracy, and national security. California Law Review, 107, 1753–1820.
- Coalition for Content Provenance and Authenticity. (2024). C2PA technical specification (Version 2.x). Linux Foundation Joint Development Foundation.
- Commissione europea. (2026). Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of Regulation (EU) 2024/1689 (C(2026) 5054 final).
- Crawford, K. (2021). Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. Yale University Press.
- Dell’Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality (Harvard Business School Working Paper No. 24-013). Harvard Business School. https://doi.org/10.2139/ssrn.4573321
- Eurostat. (2026). Use of artificial intelligence in enterprises. Statistics Explained. Publications Office of the European Union.
- Fallis, D. (2021). The epistemic threat of deepfakes. Philosophy & Technology, 34(4), 623–643. https://doi.org/10.1007/s13347-020-00419-2
- Ge, S., Mahapatra, A., Parmar, G., Zhu, J.-Y., & Huang, J.-B. (2024). On the content bias in Fréchet Video Distance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. arXiv. https://arxiv.org/abs/2404.12391
- Google DeepMind. (2025, May 20). Fuel your creativity with new generative media models and tools [Blog post].
- Groh, M., Epstein, Z., Firestone, C., & Picard, R. (2022). Deepfake detection by human crowds, machines, and machine-informed crowds. Proceedings of the National Academy of Sciences, 119(1), e2110013119. https://doi.org/10.1073/pnas.2110013119
- Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. In International Conference on Learning Representations. https://arxiv.org/abs/2009.03300
- Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., & Fleet, D. J. (2022a). Video diffusion models. Advances in Neural Information Processing Systems, 35. https://arxiv.org/abs/2204.03458
- Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., & Salimans, T. (2022b). Imagen Video: High definition video generation with diffusion models. arXiv. https://arxiv.org/abs/2210.02303
- Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. V., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., . . . Sifre, L. (2022). Training compute-optimal large language models. arXiv. https://arxiv.org/abs/2203.15556
- Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., & Liu, Z. (2024). VBench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 21807–21818). https://doi.org/10.1109/CVPR52733.2024.02060
- Interactive Advertising Bureau. (2025). 2025 video ad spend & strategy full report.
- Italia. (2025). Legge 23 settembre 2025, n. 132: Disposizioni e deleghe al Governo in materia di intelligenza artificiale. Gazzetta Ufficiale della Repubblica Italiana, Serie Generale, n. 236, 10 ottobre 2025.
- Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730
- Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate. arXiv. https://arxiv.org/abs/2509.04664
- Kang, B., Yue, Y., Lu, R., Lin, Z., Zhao, Y., Wang, K., Huang, G., & Feng, J. (2024). How far is video generation from world model: A physical law perspective. arXiv. https://arxiv.org/abs/2411.02385
- Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., & Amodei, D. (2020). Scaling laws for neural language models. arXiv. https://arxiv.org/abs/2001.08361
- Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., & Goldstein, T. (2023). A watermark for large language models. In Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 17061–17084). Proceedings of Machine Learning Research.
- Köbis, N. C., Doležalová, B., & Soraperra, I. (2021). Fooled twice: People cannot detect deepfakes but think they can. iScience, 24(11), 103364. https://doi.org/10.1016/j.isci.2021.103364
- Kuaishou Technology. (2024). Kling: Text-to-video generation model [Product announcement].
- Meng, F., Liao, J., Tan, X., Shao, W., Lu, Q., Zhang, K., Cheng, Y., & Li, D. (2024). Towards world simulator: Crafting physical commonsense-based benchmark for video generation. arXiv. https://arxiv.org/abs/2410.05363
- Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586
- OpenAI. (2024, February 15). Video generation models as world simulators [Technical report].
- Parlamento europeo & Consiglio dell’Unione europea. (2022). Regolamento (UE) 2022/2065 del Parlamento europeo e del Consiglio del 19 ottobre 2022 relativo a un mercato unico dei servizi digitali e che modifica la direttiva 2000/31/CE (Digital Services Act). Gazzetta ufficiale dell’Unione europea, L 277, 1–102.
- Parlamento europeo & Consiglio dell’Unione europea. (2024). Regolamento (UE) 2024/1689 del Parlamento europeo e del Consiglio del 13 giugno 2024 che stabilisce regole armonizzate sull’intelligenza artificiale (AI Act). Gazzetta ufficiale dell’Unione europea.
- Peebles, W., & Xie, S. (2023). Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision. https://arxiv.org/abs/2212.09748
- Rini, R. (2020). Deepfakes and the epistemic backstop. Philosophers’ Imprint, 20(24), 1–16.
- Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36. https://arxiv.org/abs/2304.15004
- Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., & Villalobos, P. (2022). Compute trends across three eras of machine learning. arXiv. https://arxiv.org/abs/2202.05924
- Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., & Taigman, Y. (2022). Make-A-Video: Text-to-video generation without text-video data. arXiv. https://arxiv.org/abs/2209.14792
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
- Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y., Zhang, F., Zhang, Y., & He, J. (2025). VBench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness. arXiv. https://arxiv.org/abs/2503.21755
- Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., & You, Y. (2024). Open-Sora: Democratizing efficient video production for all. arXiv. https://arxiv.org/abs/2412.20404
Published
2026-06-30
How to Cite
Darv, P. (2026). Synthetic realities: Generative video: creativity, epistemic limits and the verifiability of communication. AND Journal of Architecture, Cities and Architects, 49(1), 38-47. Retrieved from https://and-architettura.it/index.php/and/article/view/708