You Can’t Spell “Alternative Text” without AI -Scott Davis
About this talk
This talk explores the importance of alternative text (alt text) for images, emphasizing its role in digital accessibility for blind and low-vision users. The speaker highlights the extensive use of alt text by AI for machine learning and search engine optimization. By demonstrating how screen readers function, the speaker provides insights into how alt text effectively conveys image content when images cannot be displayed. Various examples illustrate how poor alt text can misinform users, while well-crafted text offers valuable context necessary for understanding images. The speaker also discusses legal obligations regarding accessibility and shares tips for creating effective alt text, ensuring it serves as primary content rather than an afterthought.
Full transcript
We're going to be talking about alternate text for images, and we're going to find very quickly that it's used for a variety of different reasons. One reason is for your blind and low-vision users. Did you know that India has the largest population of blind citizens in the world? So, this is an especially timely talk. Did you know that alt text is also used by AI for machine
learning? This is also a very good reason for you to be here for this talk, yes? So, let me start by telling you or showing you how a screen reader actually works. Have any of you actually ever used a screen reader before? Yes, all lovely, lovely. So, a screen reader is just what it sounds like. It's something that's going to read what's on the screen to you.
So, let me launch the screen reader right now, and I want you to look on the screen, but I also want you to listen and see if you can correspond what the computer is saying versus what you see on the screen. Here we go. Wait. >> Voice over on Chrome. You can't spell alternative text without AI Google Chrome window. An abstract visual metaphor showing a digital iris
with data streams representing machine vision image. You are currently on an image. To begin interacting with the contents of this image, press control, option, shift, down arrow. That was a lot of words, wasn't it? Some of those words were also on the screen. You can't spell alternative text without AI. But some of those words weren't on the screen. And in fact, that's the alternate text that this
describes. Voice over off. So, let's explore this idea of alt text more closely. If you see a screen like this, this is quite literally the driving instructions for this presentation. So, I can press L to toggle on my slide list to the left. I can press N to toggle on my speaker notes. But let's once again, even though there are no images on the screen, let's hear
the screen reader and how this would be presented to you. Now, this will be a memorization exercise as well, because I am going to turn on the screen reader and then turn off the So, let's do that right now. Here the screen reader comes on and I'm going to blank the screen. That's expected. We're not going to see anything, cuz this will give you the full experience
of what it's like being a screen reader user. Voice over on Chrome. Screen curtain on. You are currently on a heading level two. Go on. Keyboard shortcuts for presenting. You are currently on a selectable text. List four items. You are currently in a list. L, one of four. You are currently on a selectable list item. Toggle slide list, left side bar. N, two of four. Toggle speaker
notes, right key, three You are Toggle presenter mode, hide {slash} show {slash} space. You are move through slides. You are currently on a text element. Oh. There we go. if you've never used your ears to surf the web, that's what we want to explore here as well. But again, I can't emphasize enough that that's one use case, but not the only use case. So, we're going to
explore a lot of the different things and how this is important for AI as well. But the crux of what we want to recognize right now is that digital accessibility is sight, sound, and touch. And these are all the ways that we can surf the web and all the ways we can interact with our computer. I think many of us end up using one particular sense, don't
we? Do you listen to the source code as you're writing it as a programmer? Do you reach out and feel your source code and feel how it feels under your fingertips? No, but of course, there are braille monitors and braille keyboards. So, if you can't see the screen, these braille monitors have tiny little metal pins that pop up and down many times a second, and I'm able
to drag my finger across and read the screen through touch. So, it's an amazing world we live in right now. There are so many different ways we can consume this information. But we're going to focus specifically on alt text. This is an article from The New York Times. Alt text is for when the image can't be presented. It wants to give you an equivalent experience. So, there
is an invisible image in the middle of the screen. It didn't load, and so we're left with the alt text. A painting of a person. Maybe an image of one person and strawberries. Think for a moment. What image would you expect to appear? Did any of you guess the Mona Lisa? No. Uh, I didn't either. This was AI generated alt text. And while that alt text wasn't
wrong, it was a painting of a person, was that enough to give you an equivalent experience? Oh, and it gets worse, doesn't it? Maybe an image of one person and a strawberry? AI never hallucinates, does it? Can you see how this is even worse, isn't it? It's telling the user that there's something in the image that there clearly isn't. Now, this last bit of alt text. By
the way, these alt texts were generated Microsoft Word said a painting of a Facebook said maybe an image of a person and a strawberry. Wikipedia had this as its alt text. And as you might see, it's like, okay, Mona Lisa Leonardo da Vinci .jpeg? OH, NO. They used the file name for the alt text? Now, in this case, it was a descriptive file name, so that worked
to our benefit. What if that alt text read img001.jpeg? That too would be less informative, wouldn't it? as we're talking about this, some people hear alt text and thinks that that means optional text. Listen. We'll talk more about that. We'll give you the data back that up. But in fact, text descriptions are primary content. This is an intrinsic part of your job as a web developer. This
is an intrinsic part of what we do when we present images on the screen. It's our professional responsibility to provide alternatives for that image The first time you could use an image tag in HTML was in HTML 2. This is the HTML 2 specification. It was released in 1995. So, the first time that they mention the image element, they say the image element refers to an image.
Thank you very much. Yes, I get that. And notice in the very next sentence, they say HTML user agents, that's technical speak for the web browser. A user agent is the web browser. So, web browser may process the value of the AI attribute as an to processing the image. And this is crucial. This is crucial that the alt text be used in place of the referenced image
if it can't be loaded. If you misconfigured it so it returns a 404. If you turned off your images to preserve bandwidth. If you're driving in your car and you're using car play to hear your incoming text messages as you go along. You can see there are a variety of different use cases for this. And in this example at the bottom, image source triangle. And then the
alt is warning. The message on the screen is be sure to read these instructions and presumably this triangle image is a triangle with a red exclamation point in the middle to warn you. But if for some reason that image doesn't load, it will read warning. Be sure to follow these instructions as This was in 1995. Should I ask how many of you were born when the image
tag was released? I won't do that to you. I won't do that to you at all. Yeah, but the World Wide Web Consortium, who creates the HTML specifications, also published the web content accessibility guidelines, or WCAG. What a terrible sounding acronym, isn't it? But WCAG is what we say because web content accessibility guidelines has many many syllables. So, in WCAG, they list guidelines. And the very first
guideline they list in WCAG is for text alternatives. Guideline 1.1. They felt it was most important to refer to it first. What's nice about these guidelines is they give you success criteria. They show you what good looks like. They tell you how to succeed with these guidelines. So, here's success criterion 1.1.1, the very first success criteria. And it says all non-text content that's presented to the user
has a text that serves the equivalent purpose. So, you can see that this has been with us for quite some time. It's in the specification. It's in the web content accessibility guidelines. And there are organizations like WebAIM, Web Accessibility in Mind. They go out and survey the top 1 million websites. They do it programmatically. They crawl this and then they analyze the accessibility errors that they find.
They've been doing this for a number of years. Here's the 2026 results. Out of a million webpages, 84% of those webpages have poor color contrast. It's difficult to read. That's a lot. But look at that number two most popular accessibility error, missing alternative text for images. Even though this specification has been around since 1995, still literally, demonstrably, measurably, over half of the webpages on the internet today
lack any form of alt text at all. This is a problem. Habin Girma. I'm a big fan of Habin's. Habin Girma is deafblind. She both lacks vision and lacks hearing. So, she does all of her computer work through Braille. So, she is able to actually interact, to surf the web, to use this stuff, to send text messages, to post on Twitter. Even though she doesn't have the
visual indications or the auditory indications. So, in this tweet, she's playing with us a little bit. She's saying, "Oh boy, I'm blind. How do you describe yellow to a blind person? How do you describe red to a blind person?" And she's asking this question in the context of this image right here. So, take a moment if you would and think about what your alt text would be
for this image. Here's a great tip. When you're trying to think about alt text, think about how you might explain a photograph to a friend over the phone. If you've got the photograph, the image in front of you, and you're trying to explain it to someone verbally, how would you verbally describe what's going on in this image? You have something in mind? Yeah? Here's how Hobin described
it. I'm holding a yellow cupcake with raspberries and blueberries. What do you think? Captures what's going on in the image, doesn't it? I'm holding blueberries. The point of her tweet was a little bit ironic. She's saying, "I don't need you to explain to me what yellow is. Thank you." Right? But what I do need you to explain to me is what's going on in the image. And
I'm holding a yellow cupcake with raspberries and blueberries, I think provides the text equivalent for what's happening in that image. Hobin really appreciates that you're providing alt text in your tweets. There's some good tips for us to think about when we're writing alt text. Text description should be short and sweet and to the point. Now, they say here just more than a few words. We typically say
a sentence or two is a good target for you. So, a sentence or two, and think about the context of the Now, I just showed you an image, and she said, "This is a picture of me eating a blueberries." What if she was announcing her next book? Did I mention she's a New York Times best-selling author? Did I mention that she's the first graduate of Harvard Law
School. Her autobiography is fascinating. I highly recommend it. But what if she was announcing the publication of her second book and she said, "I'd like to share with you the cover of my next book." Does that sound like alt text to you as well? Yeah, context matters because if that's the cover of her book, she's not going to say, "Here's me eating a cupcake." Cuz that's not
the story she's trying to tell. She's trying to say, "I've got a new book coming out." You got to see that cover. here's our source code. We have an image tag. We have a source catpic.jpeg. That would be terrible alt text, Catpic.jpeg instead the alt text says, a kitten at the window. And then it goes on to describe that a robotic voice will say very quickly, a
kitten in the window image. You are interacting with the element, press control, alt, shift, delete, control Z, V, P, all of those. Yes? is one valid alt text. What if I was putting a tweet out that said, "I just bought a new cat. Please meet Tiger, my new cat." That is the context that we're looking for. And the reason I keep coming back to the context, we
are going to talk about AI quite a bit and the context is where AI needs our help the most. We're going to see that AI can describe a number of amazing things and I don't want to take away from that. Computer vision is truly amazing what it can get accomplished. But what it can't help is or the reason why we're writing this alt text in the first
place. Good. uses your alt text. It uses it quite a bit. It uses it for Google images. It uses it for page rank, for your search engine optimization. It's really amazing. It says, "Well, yeah, of course, that if you're just submitting an image and you have good alt text, that is going to help us raise the page rank of that image." Because if you think about it,
how do you search for images on the web? Text. Yes. And so, if you provide valid text for this image, that is going to improve the page ranking of that image. But, did you also know that it just improves your page rank, period? One of the best ways to improve your SEO, your search engine optimization, is to make sure that your page is well-formed. How do you
make sure that your page is well-formed? You make sure that you have alt text on all of your images. Yes, computer has computer uh Google has computer vision. It's amazing what it can do, but if you provide the context, if you provide the punchline, the meaning behind that image, that is always going to be more valuable than what the computer might try to determine on its own.
So, yes, it's here for accessibility, usability, of course. I like telling people that Google is your most frequent blind visitor to your website. Google needs your help, and you can provide it. Now, this is also required by law. Now, I'm not a lawyer, so I'm not here to tell you what you're doing wrong and claim that I'm going to arrest you, right? But, if we were in
Europe, we would be talking about the European Accessibility Act right now, the EAA. It's passed all across the EU, and it says all of your images must be closed captions. They say all of your web pages must meet these web content In the US, we have the Americans with Disabilities Act that says you must meet these web content accessibility I'm from Colorado. The state of Colorado was
the first state in the United States to pass a local digital accessibility So, we can talk about all the variety of ways where this is a legal obligation to you as well as a programmer. Section 508 is the law in the United States that dictates specifically federal agencies, what are your expectations of a federal agency when it comes to digital accessibility. And in fact, this is a
lovely website. It has a number of really good recommendations. And so, this is what we're going to walk through. But once again, good general guidelines. Your alt text should be short and sweet and to the point. It should communicate the same information as the visual content, and alt text should refer to relevant content. Am I eating yellow cupcake? Or is this the new cover of my book?
So, with all that in mind, let's go through some examples. Martin Luther King, yes? And so, at the bottom it goes through helpful and unhelpful. Um I apologize, it's a little bit low on the screen, so I will read out your alt text. So, I will be your screen reader for you. How about that? So, they say that helpful is alt text for this would be Dr.
Martin Luther King Jr. Unhelpful would be black and white photo of Dr. Martin Luther King Jr. wearing a suit and tie. You can imagine how that would be AI generated, right? There's nothing actually factually wrong with that, but is that how you would describe this image? Let's say we are on Dr. King's Wikipedia page. And this is the image that's found there on the Wikipedia page. Would
it mention that it was a black and white photo? Would it mention that he's wearing a suit and tie? Is that relevant to a Wikipedia page about Martin Luther King? Now, if I was writing an article of the best dressed civil rights activist of the 1960s, then I might mention his suit and tie. He's a sharp-looking fellow, isn't he? Right? Very well dressed. But the fact that
he's wearing a suit and tie is probably out of context for what we're trying to deal with here. So, I plugged this into Gemini. I uploaded the image and I said, "Please create alt text for the image." Are you ready for the output? A framed black and white studio portrait of Dr. Martin Luther King Jr. looking slightly to his left with a neutral expression. Dr. King is
dressed in a suit jacket, a collared shirt, and a tie with a diamond pattern. A small dark circular object is visible on the lapel of his jacket. The portrait is framed with a thick white border, which itself is framed by a gray border. The image is cropped in a vertical rectangle. This is AI, but if you'll allow me to ascribe good intentions to AI for a moment,
this is how some people write alt text. They try to say, "I want to capture every single thing about this photograph to make sure I leave nothing out, including the borders and the cropping and everything else." It is a good try. But AI is not a one-and-done exercise. AI is meant to be an ongoing iterative conversation. So, then I asked Gemini. I said, "Okay. All right. All
right. All right. Please create a concise alt text for this image that follows section 508 tips for effective alt text. I'm trying to provide it guidance for this. And it comes back and it says, "Oh, I know section 508 standards. We need to be concise and descriptive and avoid redundant phrases like image of." So, here's what I suggest, a black and white studio portrait of. Okay, all
right. It's not image of, it's not graphic of, is saying the same thing. Of Dr. Martin Luther King wearing a suit and patterned tie looking slightly off camera with a contemplative So, what I think is so ironic about this is that AI got the prompt correct and it returns these bullet items and these bullet items are all correct. It says it identifies the subject. Yes, Dr. Martin
Luther King. That's important. It avoids redundancy. It skips image of, photo of. Okay, well, no. Succinct detail. Uh no. Focuses on tent, no. Not that either, right? So, in fact, we have an AI that knows the criteria it's trying to follow. It just didn't follow the criteria that had identified. To be clear, I did that on SAS mode in Gemini. When I turned it on to thinking
mode and I said, "Please can create create a concise alt text for this image." It comes back a black and white portrait of That's the best we've seen yet, isn't Yeah. But, this is our rule of dealing with AI. You know about garbage in, garbage out, The prompts you're supplying are the garbage in at this point. It's my responsibility to provide more high-quality prompts so the AI
has an opportunity to provide more high-quality output. So, I did the same exercise in Microsoft Word. I dropped the image in Word, and I don't know if you've ever noticed at the very bottom of Word, you've got a little something that says accessibility at the bottom. If it says accessibility, good to go. You know what that means? It means you're good to go. Your accessibility is good.
That's what it means. Yeah? But if it says accessibility investigate, when you click on that, to the right, you will see the accessibility assistant. And isn't this great? It talks about color contrast. You remember that was the number one error identified by WebAIM? It goes on to talk about missing alt text, the number two error in the fix the six from WebAIM, and it goes on to
give you a bunch of things. So, when I click on missing alt text, it says, "Ah, how would you describe this object and its contents to someone who is blind or low vision? One to two sentence recommended." And there's a generate description button. What does Microsoft Word say when I ask it to generate alt text? Black and white photograph of a man wearing a suit, a white
shirt, and a pattern tie with a pin on his left lapel. The background is blurred, focusing attention on the former formal attire and upper body of the subject. Microsoft Word didn't even recognize who this was in the photo. Yeah? This is not me to continue to criticize AI, but this is me showing you in the real world how AI deals with this, and how it really is
up to us as developers to provide the context and the goodness. We can't just trust this output. We need to know what good looks like so we can help guide the AI towards better alt Going back to the New York Times article where the Mona Lisa was, they provide a number of examples. Maybe an image of fruit. AI generated. A slice of pizza sitting on top of
a white plate. Have you ever eaten pizza in your life? The AI says, "No, but I've read about it. It sounds delicious." Yes? And finally, a plate of pancakes with fruits. What if both of these images appeared in that same tweet by Haben Girma? Do you remember that the cupcake she was eating had raspberries and blueberries in it? This image right here has raspberries and blueberries in
it. So, maybe alt text in that context might be, "Here's another example of all the ways I love to eat raspberries and Pancakes. Context matters. You're trying to provide the text equivalent for why you're including the image in the first place. That's what AI struggles with. That's what you excel at, human. Maybe an image of five people standing outdoors. True. A group of people standing next to
a train. False. A group of men standing next to a helicopter. Better. That, by the way, is Facebook alt text Chrome extension in Microsoft Word. But, it might be missing the context of this. Alt text might be, let's say I'm reading this in a newspaper article and it might say, "President Joe Biden arrives in Denver for the Council of Governors on how to fight COVID-19." Does that
feel like it describes the image to you? Better than anything you saw up there so far. Memes. Memes can have alt text as well, yes? who said that? Microsoft Word said, "A picture containing text of person outdoor shore." Facebook says, "Maybe an image of two people and text that says my brain recording my good memories, my brain recording my cringe memories." It captured the text in the
image, yeah? But perhaps missed the point of the image. So, this alt text up here was written by a human and it's a nice way for you to caption memes. Top photo, person writing finger in sand as tide comes in, my brain recording my good Bottom photo, person chiseling letters into stone, my brain recording cringe Now, it may seem like it wrung every little bit of humor
out of it when you have to explain the joke to that level of detail, but again, if that image was not there, which of these alt texts would you prefer to represent the image? If you're feeding this into a large language model, which alt text would you like to be there to make sure any further work you did with this was well understood? So, we have a
number of other examples in here. One that's very interesting is an image that contains text. An image that contains text. If you have an image that does contain text, you should include your alt text word for word. So, when this image comes up, what's helpful is card with text, acquisition training for the real world, January 29th through February 9th, 1:00 p.m. to 2:00 p.m. Eastern Standard Time,
register today. It captures everything that you need to know out of that image. What's unhelpful? Event details with a registration What are you left to asking? What are the details? What's the event? How do I register? I experience this. I've been coming to Bengaluru now for 18 years. When you're looking at train tables and bus schedules and things like that, it's not uncommon at all for them
to be presented in a JPEG. And when I go to translate Hindi to English, all of the text on the web page turns into English, and the text captured in that JPEG remains the same. So, text in image can be particularly tricky, but luckily we have alt text. If you provide the text in the image in your alt text, you've provided an Logos are much the same
way. It's very common for someone to say logo or GSA's logo. So, when we have a logo like this, logos are never decorative, so they require alt text. Describe any significant symbols, including any text in the logo word for word. So, this is GSA section 508.gov, build by build, be accessible. That's the alt text they provide. Yeah? What's unhelpful? Logo. So, it did hit on an important
concept here that we do have this notion of decorative images as well. We have horizontal lines across web pages separating sections. We have vertical lines. We have decorative elements, all those things. Believe it or not, you can flag that as decorative so it won't be read to the screen reader. Imagine I'm your screen reader right Heading level two, course completion. Image. Image of a woman typing on
a laptop. Now that you've completed this course, you should be able to do the following. Was the image of a woman of a woman typing on a laptop important to the context of this? Truly decorative. And what's interesting is if it is decorative, you don't just leave out the attribute. You in fact say alt equals quote quote, empty string. That is the signal to the screen reader
that I acknowledge that this is decorative. I acknowledge there's nothing you should say, so alt equals quote quote is the way you say to the screen reader, don't say anything at But what's important is that you don't leave off the alt attribute because some screen readers might go in and then read the file name in its place. So, if I'm the screen reader and I say, heading
level two, course completion, File name, img001.jpeg. Knowing where your images are decorative can be very powerful as well. All of this time we've been talking about how you convince the screen reader to say something important, you can also say there's nothing important to see here, move along. Move along. It's so common for us to have background Do you think background images are the important part of the
story? It's decorative as well. So, if the screen reader said, "Let's start with a brief overview of what Section 508 is and what it involves. Image, keyboard, white keys." It doesn't contribute to the conversation, does it? So, we know that decorative images are there and available for us as well. They can give visual insight, but if they're they're there. What time does this session end? Are we
at time right now? Are we getting close? 5 minutes? No, we have 10. Well, how how much time left? Well, wait. 20 minutes? Okay, very good. Thank you. when we're dealing with controls, dealing with controls on screen, this is a really crucial aspect of navigating your web pages. All of a sudden, we're not explaining what's in the image, but we're explaining why this image is there and
what it does if you interact with it. So, we have two images right here, previous and next. It could be implemented a number of different ways. Those could be a link surrounding that image. So, as you click on the previous button, it goes to slide 37. As you click on the next button, it goes to slide 39. Yeah? And so, helpful would be a screen reader reading
the links next and previous and ignoring the images all together. The important part of this is supplying the link in the behavior. It might be less important that it was a blue arrow with white text on it indicating if you press this, you will go to the previous slide. Yeah? If you do put alt text on there next arrow button and previous arrow button, do a good
job of explaining exactly what those images are and what they're for. If we provided this alt text and just it simply said, "Blue arrow with text." That would probably not be as helpful as it could be. we've covered a lot of ground here. And again, my point is not for us to make fun of all the mistakes that AI makes. That was never my intention at all.
In fact, I use AI quite a bit for these kinds of things, but it has to be an active engagement with your LLM, and you need to bring what good looks like to the LLM through your prompts and through the iterative conversation you have with these things. So, you don't want your alt text to be redundant. So, if it's already being announced somewhere, it's fine to make
that You never want file names in your alt text, although ironically, if you eliminate the alt text attribute all the file name is often what's read. So, developers then begin thinking, "Oh, well, if this is default behavior, what would be better than me making that behavior explicit?" No, that's not what we want either. We've talked about internationalization. Again, text on the page can be translated from Hindi
to Tamil to English. Pixels in a JPEG can't be translated nearly as well. And descriptive clutter. And this is where judgment comes in. This is where you need to use your common sense to hear is this conveying what this image is for? Is this saying the right thing? Or is this saying someone in a blue shirt and a blue sports jacket and tan pants with shoes was
up here talking about alt text. No lies detected, but that is There are times when you truly feel like you've got to pack information in there. If this is an image of a chart or a you might be compelled to say, "All right, sales figures for 2025, in January we sold this, in February we sold this, in March we sold this, in April we" Yeah? Yeah? Chances
are you don't want that information in the alt text. Chances are you want that kind of information in a data table just in the page for every everyone to see. You can include that chart, but if you include the specifics of that in a data table and text in there, then you might even be able to flag that chart as decorative if all of the data is
there for people to read. Or if that chart is saying "Sales were going down through August and then spiked for the rest of the year." That's great alt text. That's conveying what that chart is trying to tell you. And then by providing a data table beneath it, you then give everyone the opportunity to evaluate that alt text to make sure that they agree with you as well.
So, there are some great attributes we can use. Aria described by a detailed summary set of elements. Um detailed summary can even be configured so it's expandable and collapsible. So, you might have a chart telling that sales were going down until August and then they spiked, and then have a data table that's collapsed. And so then people could actually expand it if they wanted to learn more
about it, but it would be visually hidden if they said, "Nah, I saw the chart. I got the point." But this is what I want you to leave you with. And I hope every single person who you hear today talking about AI stresses this. This is a human AI collaboration. You must be just as engaged with your AI as you would be with a pair programming partner.
Or a junior developer you're trying to explain your code base to. So, in this collaboration we have to be specific. You can't just say, "Hey, make me some alt text." You might say, "Here is the Wikipedia page for Martin Luther King Jr. Can you please provide an alt text for this that conforms to Section 508 standards?" I have provided so much rich context for that AI that
that AI will have a much higher opportunity to get it right. And that's hard. And that's why we're talking about using AI to generate alt Because as developers, I'm a developer myself, sometimes we struggle with naming variables, don't we? foo baz, x If you're having a hard time naming your variables, maybe you have two challenges. Maybe you could feed your code into an AI, an LLM, and
say, "Could you help me better name my variables? Could you help me create more readable code? Code that's easier to debug? Code that's easier to share with my colleagues?" I would worry less about hallucination then, because I'm providing it code and asking it to change the names of my variables in there. That's the level of You can't say, "Here's my code, make it better." I wish we
could. I wish we could. But if you said, I've gotten feedback that my variable names aren't descriptive enough. Can you help me with that?" is a much different prompt. Keep it concise is talking about the alt text, not your prompts. Some of those prompts I gave you are actual prompts that I've provided to uh LLMs. It's ironic that sometimes I will spend two paragraphs in a prompt
and press enter and get one sentence back. That seems like a mismatch in the level of effort versus the quality of the output, but in fact, that's exactly the level of input I needed to provide to get the quality output that I was Be skeptical. Don't assume that the AI is getting it wrong. Hey, everyone, watch this. Watch it what it says this time, right? Every time
an AI gets alt text wrong, I blame myself. What could I have done differently in that prompt? What could I have done to give it more context? What could I have done to ensure better output? But when it does come back, remember this is not one shot conversations with AI. Very frequently, I'll come back and say, "That's not what I was looking for at all. Let me
try again." Or sometimes I'll just say, "You try again." No, that wasn't it either. Try again. I don't get very good output when I do that. Just try again. And I say, "Ah, we're not interested in what Martin Luther King is wearing in this image. Can you try again?" the fact that this is a black and white image isn't important in context. Can "There is no strawberry
in the photo in the in the painting of the Mona Lisa. Please try again." But at the end, make it your own. At the end, I'll often times get output from an AI that's 80% of the way there, and that's all I needed. I needed it enough to get me close, but then as the human in the loop, I was the one that took it the rest
of the way and completed it. I've had the same experience that I showed you here where AI will come back and say a black and Jr. I'm like, "You know what? I'm going to strip off a black and white photo of and then I have what I need." I was worried that maybe the title of this talk would be too clever by half. a Well, there's an
A in alternative text and an I in alternative text, right? Yes. Yes. But, the flip side of that is AI can't do it without you, either. And that's what gets lost in this So many times people look at AI and they say, "Hey dude, write my term paper for me." That's not a responsible use of AI, And can you imagine what kind of quality term paper you
might get out of ChatGPT with a prompt like that? on American history." this standard has been in place since the beginning of images on the web. But now in this age of AI good quality alt text is more important than ever because in this human-AI collaboration, alt text is how you provide context. Alt text is how you provide the information you need. You know? Do you remember
when you would try to log into a website, it would provide you with a CAPTCHA? And it would provide you with six images and say, "Select all of the stop signs in these images. Select all the crosswalks in these images. Select all the school buses in these images, right?" What were you doing when you were doing They claim that you were proving that you were a human.
But do you know what else you were doing? You were training the AI. You were the human in the loop. And so as we're writing our alt text, I encourage you to continue to be the human in the loop. Thank you so much for your time and attention. I really do appreciate it.
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59