Mahendranagar, 18 August
The flurry of images generated by artificial intelligence (ai) feels like the product of a thoroughly modern tool. In fact, computers have been at the easel for decades. In the early 1970s Harold Cohen, an artist, taught one to draw using an early ai system. “Aaron” could instruct a robot to sketch black-and-white shapes on paper; within a decade Cohen had taught Aaron to draw human figures.
Today “generative AI” models put brush to virtual paper: publicly available apps, such as Midjourney and OpenAI’s dall-e, create images in seconds based on text prompts. The final products often dupe humans. In March ai-generated images of Donald Trump being handcuffed by police went viral online. And image generators are improving fast. How do they work—and how are they refining their craft?
Generative-ai models are a type of deep learning, a software technique that uses layers of interconnected nodes that loosely mimic the structure of the human brain. The models behind image-generators are trained on enormous datasets: laion-5b, the largest publicly available one, contains 5.85bn tagged images. Datasets are often scraped from the internet, including from social-media platforms, stock-photo libraries, and shopping websites.
The most sophisticated image generators frequently employ a diffusion model, a kind of generative artificial intelligence. visuals in the dataset are given distorted visual "noise" until the visuals are totally covered, giving the impression that analog television signals are still interrupted by static. The model can create an image that is comparable to the original by learning how to undo the mess. It starts to condense, classify, and store this knowledge in a mathematical pocket of code known as the "latent space" as it gets better at identifying clusters of pixels that correspond to specific visual notions.
Consider asking an app to make an image of a hippopotamus. In order to produce a realistic image of the mammal, a model should be able to sample from its latent space after learning which kinds of pixel arrangements correspond to the phrase "hippopotamus" (see image, left). Adding more information to the prompt—for instance, "an oil painting from the Renaissance of a green hippopotamus, somewhere along the Nile" (see image, right)—requires the model to locate and properly combine additional layers of visual detail, such as image style, texture, color, and location.
The responses to complicated prompts can be erratic, particularly if the prompt is not clearly phrased or the scene it describes is not well represented in the training dataset. Even seemingly simple fare can trip models up. Human hands are often depicted with missing or extra fingers, or proportions that appear to bend the rules of physics. Because hands are usually less prominent than faces in photographs, there are smaller datasets for ai models to hone their technique on. Dodgy facial symmetry—especially inconsistencies in colour and shape between eyes, teeth and ears—is another sign of a machine’s work. And image generators struggle with text, often creating non-existent letters or imaginary words.
Developers can help models to learn from their mistakes by refining the datasets that they are learning from or by tweaking algorithms. Midjourney was recently updated to improve the way it generates hands. Rapid improvements mean that telling an ai-generated image from a real photograph or painting may soon become impossible.
.gif)