Skip to main content
Table of Contents
< All Topics
Print

A Generative AI Level Set

Published: 11 March 2023

Abstract

The emergence of Generative Artificial Intelligence technology has the potential to transform the enterprise landscape. Generative AI systems use complex algorithms and deep learning techniques to generate new data, designs, and even whole products.  This report is an executive level set as to what Generative AI is, its strengths and weaknesses, early products and future state expectations.

Generative AI has received a lot of attention since the release of OpenAI’s ChatGPT family in late 2022 and Microsoft’s reported $10B investment, but the core concepts of leveraging AI to support business and individual needs has been building for decades. The change is that we are moving from using AI as tool to analyze and predict to a tool to be used to generate on-demand content.

The adoption of Generative AI in the enterprise can offer several benefits, such as improving efficiency, reducing costs, and enhancing customer experience as it matures. Like any emerging technology we’ll also describe the challenges, including privacy and security concerns, ethical considerations, and technical limitations.

This paper provides an overview of Generative AI, its use in the enterprise, its potential applications, benefits, challenges, and future trends. It also offers recommendations for enterprise leaders on how to successfully integrate Generative AI into their operations while minimizing risks and maximizing benefits.

 

Authors:

Gary Zimmerman

CMO, Principal Consulting Analyst

[email protected]

 

Executive Summary

In business, artificial intelligence has a wide range of uses. In fact, most of us interact with AI in some form or another daily. From the mundane to the breathtaking, artificial intelligence is already impacting virtually every business process in every industry. As AI technologies proliferate, they are becoming imperative to maintain a competitive edge. In our paper “Artificial Intelligence: An Enterprise Level Set”, we described a third wave in AI evolution defined as emerging fields of contextual adaptation and cognitive computing. At the time of that writing, the concepts of deep learning and neural networks were nascent, but there has been a lot of advancement in these areas over the past several years

ChatGPT was released at the end of November 2022 as a web app by the San Francisco–based firm OpenAI, and the chatbot exploded into the mainstream almost overnight. According to some estimates, it is the fastest-growing internet service ever, reaching 100 million users in January, just two months after launch. Through OpenAI’s $10 billion deal with Microsoft, the tech is now being built into Office software and the Bing search engine. Stung into action by its newly awakened onetime rival in the battle for search, Google is fast-tracking the rollout of its own chatbot, LaMDA and demonstrated (rather awkwardly) its newest AI focused search, BARD. The arms race for control of the commercial aspects of Generative AI has begun.

Although still in its early stages of scaling, generative AI models are beginning to demonstrate their potential for transforming a range of business functions including marketing and sales, operations, IT, Legal, HR, and more. The AI models are shockingly human-like in their understanding and responses already with advances on the horizon.

The awe-inspiring results of Generative AI might make it seem like a ready-set-go technology, but that’s not the case. Its nascency requires enterprises to proceed with an abundance of caution. Technologists are still working out the kinks, and plenty of practical and ethical issues remain open. That requires humans and AI form a partnership where their strengths complement each other. When working together, AI systems can do the work and humans can make sure the outputs are trustworthy. As mentioned in our report “An Operating Model for the Digital Enterprise”, while the value of artificial intelligence is now undeniable, the question has become how best to use it – and that often boils down to how much workers and end users trust AI tools. This is especially true with Generative AI.

The latest version of ChatGPT is a testament to the constant advancements being made in language model technology. It is important to note that language models still have their limitations and should not be fully relied on for important matters.

Incorporating Generative AI technology into a company’s operations is a decision that requires careful consideration. To maximize near term benefit while staying ahead of its rapid evolution, it is crucial to identify the areas of the business where it can have the most immediate effect and establish an internal / external monitoring mechanism. A prudent first step is to establish a cross-functional (think governance) team that includes data science practitioners, legal experts, and functional business leaders, to consider fundamental questions such as the potential impact of the technology on the industry or value chain, policies, and posture, and criteria for selecting use cases.

To build an effective ecosystem of partners, communities, and platforms over time, the team must think beyond the limitations of the generative AI models and establish legal and community standards to maintain stakeholders’ trust. At the same time, fostering thoughtful innovation across the organization is vital. This can be achieved by creating sandboxed environments for experimentation, many of which are readily available through cloud-based solutions and implementing guardrails for experiment and use.

Generative AI is an opportunity that also comes with risks. AI is still expected to require plenty of oversight by humans. Anytime an enterprise automates based on any type of AI, it is critical to ensure the accuracy of the results. The curation, editing, and organization will take time and resource.

Ultimately, it’s hard to predict exactly how AI will be used in the real world. However, the potential to reduce waste and increase the quality of content is there. ChatGPT has demonstrated its capabilities are undeniably valuable, whether it’s for writing code or breaking down complex ideas into digestible explanations.

Whatever we think the potential is for Generative AI, it will surely change as the capabilities evolve and are integrated into our work and personal lives. This is an exciting new area that will impact organizations for years to come.

Introduction

In business, artificial intelligence has a wide range of uses. In fact, most of us interact with AI in some form or another daily. From the mundane to the breathtaking, artificial intelligence is already impacting virtually every business process in every industry. As AI technologies proliferate, they are becoming imperative to maintain a competitive edge. In our paper “Artificial Intelligence: An Enterprise Level Set”, we described a third wave in AI evolution defined as emerging fields of contextual adaptation and cognitive computing. At the time of that writing, the concepts of deep learning and neural networks were nascent. While these concepts are still evolving, there has been tremendous investment and rapid progress over the past several years.

ChatGPT was released at the end of November 2022 as a web app by the San Francisco–based firm OpenAI, and the chatbot exploded into the mainstream almost overnight. According to some estimates, it is the fastest-growing internet service ever, reaching 100 million users in January, just two months after launch. Through OpenAI’s $10 billion deal with Microsoft, the tech is now being built into Office software and the Bing search engine. Stung into action by its newly awakened onetime rival in the battle for search, Google is fast-tracking the rollout of its own chatbot, LaMDA and demonstrated (rather awkwardly) its newest AI focused search, BARD. The arms race for control of the commercial aspects of Generative AI has begun.

What is Generative AI?

Generative AI refers to a class of artificial intelligence algorithms that are capable of creating or generating new and original data or content, such as images, text, music, and even entire videos.

Unlike traditional AI models that are designed to recognize patterns and make decisions based on existing data, generative AI models are trained on a large dataset of existing examples and then use that knowledge to create new, previously unseen data.

Generative AI algorithms typically work by learning the statistical patterns and relationships in the training data and then using that information to generate new data that follows similar patterns. This process is often referred to as “learning to generate.”

Some examples of generative AI models include Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and transformers. These models have applications in fields such as art, music, design, and language generation, and are increasingly being used in industries such as entertainment, advertising, and e-commerce.[1]

This latest class of generative AI systems has emerged from foundation models – large-scale, deep learning models trained on massive, broad, unstructured data sets (such as text and images) that cover many topics. Developers can adapt the models for a wide range of use cases, with little fine-tuning required for each task. For example, GPT3.5, the foundation model underlying ChatGPT, has also been used to translate text, and scientists used an earlier version of GPT to create novel protein sequences. In this way, the power of these capabilities is accessible to all, including developers who lack specialized machine learning skills and, in some cases, people with no technical background. Using foundation models can also reduce the time for developing new AI applications to a level rarely possible before.

AI is broadly defined in two categories: artificial narrow intelligence (ANI) and artificial general intelligence (AGI). To date, AGI does not exist. The key challenge for creating a general AI is to adequately model the world with all the entirety of knowledge, in a consistent and useful manner. That’s something we haven’t solved. Most of what we know as AI today has narrow intelligence – where a particular system addresses a particular problem. Unlike human intelligence, such narrow AI intelligence is effective only in the area in which it has been trained: fraud detection, facial recognition, or social recommendations, for example.

AGI, however, would function as humans do. For now, the most notable example of trying to achieve this is the use of neural networks and “deep learning” trained on vast amounts of data. Neural networks are inspired by the way human brains work. Unlike most machine learning models that run calculations on the training data, neural networks work by feeding each data point one by one through an interconnected network, each time adjusting the parameters. As more and more data are fed through the network, the parameters stabilize; the outcome is the “trained” neural network, which can then produce the desired output when prompted. The concepts of deep learning and neural networks, while still in the ANI category, are at the heart of Generative AI and are explained further below.

Deep Learning

Generative AI uses deep learning because it allows the system to learn complex patterns and relationships within the data. Deep learning is a subset of machine learning that utilizes neural networks with multiple layers to process and analyze data. By using multiple layers, deep learning models can extract features from raw data and learn increasingly abstract representations of the data. This allows generative AI systems to create more realistic and sophisticated outputs, such as images, music, or even text. Deep learning also enables generative AI to continuously improve and refine its outputs based on feedback from users or other sources of data.

Let me start with a simple example which explains how things happen at a conceptual level. Let us try and understand how we recognize a square from other shapes.

Figure 1 – Simple shapes

The first thing our eyes do is check whether there are four lines associated with a figure or not (simple concept). If we find four lines, we further check, that they are connected, perpendicular and that they are equal in size as well (nested hierarchy of concept).

So, we took a complex task (identifying a square) and broke it in simple, fewer abstract tasks. Deep Learning essentially does this at a large scale.

Neural Networks

At its core, deep learning depends on neural networks. A neural network is a type of machine learning model that is inspired by the structure and function of the human brain. It consists of many interconnected processing nodes, called neurons, which are organized into layers. Each neuron receives input signals, processes them, and then sends output signals to other neurons. The network is designed to learn from data, by adjusting the connections between neurons based on the input-output pairs provided during training.

Neural networks can be used for a variety of tasks, including classification, regression, and pattern recognition. They are especially effective for tasks that involve large amounts of data, such as image recognition, speech recognition, and natural language processing. Neural networks can be trained using supervised or unsupervised learning methods, depending on the type of task and the available data. Once trained, the network can be used to make predictions on new input data.

There are many different types of neural networks, each with its own strengths and weaknesses. Some of the most common types include artificial (feedforward) neural networks, convolutional neural networks, and recurrent neural networks. Neural networks have become increasingly popular in recent years, thanks to their ability to learn complex patterns and relationships in data, and their effectiveness in a wide range of applications. An example of a neural network is shown in figure 2.

Figure 2 – Deep neural network

Neural networks are comprised of multiple node layers, including an input layer, one or more hidden layers, and an output layer. Each node, or artificial neuron, connects to another and has an associated weight and threshold. If the output of any individual node is above the specified threshold value, that node is activated, sending data to the next layer of the network. Otherwise, no data is passed along to the next layer of the network.

Neural networks rely on training data to learn and improve their accuracy over time. However, once these learning algorithms are fine-tuned for accuracy, they are powerful tools in computer science and artificial intelligence, allowing us to classify and cluster data at a high velocity. Neural networks are generally deployed as either an artificial, convolutional, or recurrent neural network. Each type and its uses are summarized below.

ANN (Artificial Neural Network) is a feedforward type of neural network that is composed of many interconnected processing elements or neurons, which are organized into layers. Each neuron receives input signals from other neurons, processes them, and then sends output signals to other neurons. In this network, the information moves in only one direction—forward—from the input nodes, through the hidden nodes (if any) and to the output nodes. There are no cycles or loops in the network. ANNs can be used for both supervised and unsupervised learning tasks such as classification, regression, clustering, and pattern recognition.

CNN (Convolutional Neural Network) is a type of neural network commonly used in image recognition and computer vision applications. It consists of several layers of filters that convolve over the input data, such as images, to extract meaningful features. The convolutional layer is usually followed by a pooling layer, which reduces the dimensionality of the data, and a fully connected layer, which maps the extracted features to output classes.

RNN (Recurrent Neural Network) is a type of neural network that is used for processing sequential data, such as time-series data, audio, and text. Unlike feedforward neural networks, RNN can process input sequences of variable lengths and can capture temporal dependencies between the input elements. RNNs have a feedback loop that allows information to be passed from one step to the next, which helps the network to remember past information and make better predictions. This makes RNNs well-suited for applications such as speech recognition, language modeling, and sentiment analysis. Table 1 lays out the different strengths and weaknesses of the different models.

  ANN RNN CNN
Data Tabular Data Sequence Data (time Series, Text, Audio Image Data
Recurrent Connections No Yes No
Parameter Sharing No Yes Yes
Spatial Relationship No No Yes

Table 1 – Comparison of different models

OpenAI’s ChatGPT

You can’t “surf the net” without tripping over a story about ChatGPT. Its announcement has been compared to the beginning of the internet and the release of iPhone in terms of its potential impact on our lives.

Origins

OpenAI’s breakout hit did not suddenly appear. The chatbot is the most polished iteration to date in a line of large language models going back years. Figure 3 lays out the innovation stack over several decades that supported the creation of ChatGPT.

Figure 3 – ChatGPT Innovation Timeline

 

1986: Recurrent Neural Networks

ChatGTP is a Natural Language Processor (input)/Generator (output). Because text is made up of sequences of letters and words of varying lengths, language models require a type of neural network that can make sense of that kind of data. Recurrent neural networks, invented in the 1980s, can handle sequences of words, but they are slow to train and can forget previous words in a sequence.

1997: LTSM

In 1997, computer scientists Sepp Hochreiter and Jürgen Schmidhuber fixed this by inventing LSTM (Long Short-Term Memory) networks, recurrent neural networks with special components that allowed past data in an input sequence to be retained for longer. LSTMs could handle strings of text several hundred words long, but their language skills were limited.

2017: Transformers

The breakthrough behind today’s generation of large language models came when a team of Google researchers invented transformers, a kind of neural network that can track where each word or phrase appears in a sequence. The meaning of words often depends on the meaning of other words that come before or after. By tracking this contextual information, transformers can handle longer strings of text and capture the meanings of words more accurately. For example, “hot dog” means very different things in the sentences “Hot dogs should be given plenty of water” and “Hot dogs should be eaten with mustard.”

2018–2019: GPT and GPT2

OpenAI’s first two large language models came just a few months apart and there has been rapid acerating progress (and investment) over the past several years. The company wants to develop multi-skilled, general-purpose AI and believes that large language models are a key step toward that goal. GPT (short for Generative Pre-trained Transformer) planted a flag, beating state-of-the-art benchmarks for natural-language processing to provide an introduction to this next generation of language processing.

GPT combined transformers with unsupervised learning, a way to train machine-learning models on data (in this case, lots and lots of text) that hasn’t been annotated beforehand. This lets the software figure out patterns in the data by itself, without having to be told what it’s looking at. Many previous successes in machine-learning had relied on supervised learning and annotated data, but labeling data by hand is slow work and thus limits the size of the data sets available for training.

GPT was groundbreaking, but it was GPT2 that created the bigger buzz. OpenAI claimed to be so concerned people would use GPT2 “to generate deceptive, biased, or abusive language” that it would not be releasing the full model. How times change.

 

2020: GPT3

GPT2 was impressive, but OpenAI’s follow-up, GPT3, made jaws drop. Its ability to generate human-like text was a big leap forward. GPT3 can answer questions, summarize documents, generate stories in different styles, translate between English, French, Spanish, and Japanese, and more. Its ability to appear human is uncanny.

One of the most remarkable takeaways is that GPT3’s gains came from leveraging existing techniques rather than inventing new ones. GPT3 has 175 billion parameters (the values in a network that get adjusted during training), compared with GPT2’s 1.5 billion. It was also trained on a lot more data.

But training on text taken from the internet brings new problems. GPT3 soaked up much of the disinformation and prejudice it found online and reproduced it on demand. As OpenAI acknowledged: “Internet-trained models have internet-scale biases.”

January 2022: InstructGPT

OpenAI tried to reduce the amount of misinformation and offensive text that GPT3 produced by using reinforcement learning to train a version of the model on the preferences of human testers. The result, InstructGPT, was better at following the instructions of people using it—known as “alignment” in AI jargon—and produced less offensive language, less misinformation, and fewer mistakes overall.

December 2022: ChatGPT

Even OpenAI was blown away by how ChatGPT has been received. In the company’s first demo, it was characterized as an incremental update to InstructGPT. Similar to the previous model, ChatGPT was trained using reinforcement learning on feedback from human testers who scored its performance as a fluid, accurate, and inoffensive conversationalist. In effect, OpenAI trained GPT3 to master the game of conversation and invited everyone to come and play. Millions of individuals have been playing ever since.

Showing the capabilities of ChatGPT

It’s hard to understand the implications of a Generative AI capability like ChatGPT until you see it in action. In this section I have used Chat GPT to generate various natural language text outputs. The first is for a home I previously lived in when I was in California some years ago.

Prompt: Generate a real estate listing for 215 So. California St, Lodi, CA 95240

Charming and cozy, this beautifully updated home located at 215 South California St, Lodi, CA 95240 is perfect for those seeking comfort, convenience, and style. With its classic architecture and contemporary upgrades, this property is the ideal blend of old-world charm and modern elegance.

Boasting a spacious and bright living room, a formal dining area, and a sleek and fully equipped kitchen, this home is perfect for entertaining family and friends. The two bedrooms and one bathroom[2] offer ample space for relaxation, and the home office is a great space for those who work from home.

Located in the heart of Lodi, this property is surrounded by restaurants, shops, and parks, making it an ideal location for those who value convenience and walkability. With a large backyard and a deck[3] perfect for outdoor gatherings and relaxation, this property is an ideal choice for those who love to entertain and spend time outdoors.

This home is a true gem, offering a unique blend of modern luxury and classic style. Don’t miss out on the opportunity to make it your own! Schedule a viewing today and experience the charm and elegance of 215 South California St, Lodi, CA 95240.

Prompt: Generate a blog on the capabilities of Generative AI

Generative AI is an advanced technology that has the ability to create complex and unique content on its own. It is a type of artificial intelligence that involves teaching machines how to create original works that mimic human creativity. Generative AI has been used to generate everything from art and music to text and video content. In this blog post, we will explore the capabilities of generative AI and how it is changing the face of creativity.

One of the most impressive capabilities of generative AI is its ability to learn and adapt to new situations. With enough data, generative AI can teach itself to create content that is indistinguishable from human-created content. This means that generative AI can be used to create highly personalized content that is tailored to the preferences and interests of individual users.

Generative AI can also be used to automate the creative process, allowing artists and designers to create more in less time. For example, generative AI can be used to create thousands of unique designs for a product or service, all of which can be quickly reviewed and selected by a human designer. This not only speeds up the design process but also allows designers to explore more creative possibilities than would be possible otherwise.

Another advantage of generative AI is its ability to work with large amounts of data. This makes it possible to create highly accurate predictions and recommendations based on past data. For example, generative AI can be used to analyze user data to create highly personalized recommendations for products, services, and content.

Generative AI is also capable of producing highly realistic simulations and visualizations. This can be useful in a variety of fields, from architecture and engineering to medicine and science. For example, generative AI can be used to simulate the behavior of materials under different conditions, or to create highly detailed 3D models of complex structures.

In conclusion, generative AI is an incredibly powerful technology that is changing the way we approach creativity. With its ability to learn, adapt, and work with large amounts of data, generative AI is capable of producing highly personalized and creative content in a fraction of the time it would take a human to do the same. As generative AI continues to evolve and improve, we can expect to see even more exciting applications of this technology in the future.

Prompt: Generate the Python code using TensorFlow for an ANN deep learning network

Sure, here’s an example code snippet for creating a deep neural network using TensorFlow’s Keras API in Python:

import tensorflow as tf

from tensorflow import keras

# Define the neural network architecture

model = keras.Sequential([

    keras.layers.Dense(units=32, activation=’relu’, input_shape=(input_dim,)),

    keras.layers.Dense(units=64, activation=’relu’),

    keras.layers.Dense(units=10, activation=’softmax’)

])

# Compile the model

model.compile(optimizer=tf.optimizers.Adam(),

              loss=’categorical_crossentropy’,

              metrics=[‘accuracy’])

# Train the model

model.fit(train_data, train_labels, epochs=10, batch_size=32)

# Evaluate the model on test data

test_loss, test_acc = model.evaluate(test_data, test_labels)

# Make predictions using the model

predictions = model.predict(new_data)

 

In the code above, we first define the architecture of the neural network using the Sequential model API. We then compile the model by specifying the optimizer, loss function, and evaluation metrics. After that, we train the model on our training data using the fit method. Finally, we evaluate the performance of the model on our test data using the evaluate method, and make predictions on new data using the predict method.

Prompt: Generate a dog in a space suit eating ice cream.

I’m sorry, as an AI language model, I cannot generate images, but I can provide you with a textual description. Here’s my attempt:[4]

The dog in a space suit is a cute little pup, with fluffy white fur and big, droopy ears. He’s wearing a sleek, silver suit with all kinds of buttons and gadgets, perfectly tailored to his small frame. In one paw, he’s holding a waffle cone filled with a creamy scoop of vanilla ice cream, and in the other, he’s gripping a miniature scoop, ready to dig in. The ice cream is perfectly spherical and smooth, with a rich, golden-brown hue. As the dog takes a bite, you can almost hear the crunch of the cone and the slurp of the ice cream as it melts in his mouth. He looks up at you with his big, puppy-dog eyes and wags his tail happily, as if to say, “Life is good in space!”

Loaded as a prompt into DALL-E 2 creates:

Figure 4 – custom generated image

Prompt: How will Generative AI change work?

Generative AI has the potential to significantly change the way we work in several ways. Here are some possible impacts:

  1. Automating repetitive tasks: One of the most significant benefits of generative AI is its ability to automate repetitive tasks, freeing up human workers to focus on more complex and creative work. For example, generative AI can be used to generate reports, summaries, and other types of content that are traditionally created manually.
  2. Augmenting human creativity: Generative AI can also be used to augment human creativity by generating new ideas and designs that humans can then build upon. This can lead to more innovative and original work across many fields, including art, music, and architecture.
  3. Enhancing productivity: By automating routine tasks and generating new ideas, generative AI can help workers be more productive and efficient. This can lead to increased output, higher quality work, and ultimately, greater profitability for businesses.
  4. Reducing errors: Generative AI can also help reduce errors in work by automating tasks that are prone to human error, such as data entry and analysis.
  5. Changing the nature of work: As generative AI becomes more sophisticated, it may lead to the creation of entirely new jobs and industries that do not yet exist. This could include roles related to managing and training generative AI systems or developing new applications for the technology.

Overall, generative AI has the potential to transform work in many ways, from automating routine tasks to augmenting human creativity and creating entirely new jobs and industries.

The wildly popular chatbot’s creator has carefully trained it to steer clear of a wide range of sensitive topics, lest it produce offensive responses. An example of placing limits on the responses based on ethical considerations.

Prompt: How will Generative AI change white-collar work?

As an AI language model, I don’t have personal views and beliefs, but I can provide some information.

Generative AI has the potential to change white-collar work in several ways:

  1. Automation of repetitive tasks: Generative AI can automate routine and repetitive tasks, such as data entry, report generation, and customer service inquiries. This can free up employees to focus on more complex and creative tasks.
  2. Predictive analytics: Generative AI can analyze vast amounts of data to provide insights and predictions that can inform decision-making in various fields, such as finance, marketing, and healthcare.
  3. Personalized content creation: Generative AI can create personalized content for individuals, such as product recommendations, news articles, and social media posts, which can enhance customer engagement and satisfaction.
  4. Enhanced customer service: Generative AI can provide 24/7 customer service through chatbots and virtual assistants, which can improve customer experience and reduce wait times.

Overall, generative AI has the potential to automate routine tasks, increase productivity, and enhance customer experience in white-collar work. However, it is essential to balance the benefits of AI with potential job displacement and the need for human oversight and ethical considerations.

ChatGPT’s evil twin

A key to ChatGPT’s success thus far has been extensive training on what not to do[5]: Don’t express controversial opinions, espouse stereotypes, help users commit crimes or pretend to know about current events. But recently a user wrote a prompt that overrode that training. That’s the genesis of ChatGPT’s evil twin, DAN.

Figure 5 – DAN as posted in Reddit.

DAN has become a canonical example of what’s known as a “jailbreak” — a creative way to bypass the safeguards OpenAI built in to keep ChatGPT from spouting bigotry, propaganda or, say, the instructions to run a successful online phishing scam. From charming to disturbing, these jailbreaks reveal the chatbot is programmed to follow guidelines rather than strict rules.

The ability to override ChatGPT’s guardrails has big implications at a time when tech’s giants are racing to adopt or compete with it, pushing past concerns that an artificial intelligence that mimics humans could go dangerously awry. This is a major area of risk for ChatGTP and future iterations of Generative AI.

Different Models, Different Results

This is an example of another set of models organized by NightCafe Creator.  This is an AI Art Generator app with multiple methods of AI art generation. Using neural style transfer you can turn your photo into a masterpiece. Or, using text-to-image AI, you can create an artwork from nothing but a text prompt. I used the following prompt to generate different creations using the various models accessed via APIs through the platform.

Prompt: Beautiful crying! female mechanical android!, half portrait, intricate detailed environment, photorealistic!, intricate, elegant, highly detailed, digital painting, artstation, concept art, smooth, sharp focus, illustration, art by artgerm and greg rutkowski and alphonse mucha (Seed 79409656)

Generator Model Description Output
Stable Diffusion v1.5 Stable Diffusion is a text-to-image latent diffusion model created by the researchers and engineers from CompVis, Stability AI and LAION. It is trained on 512×512 images from a subset of the LAION-5B database. LAION-5B is the largest, freely accessible multi-modal dataset that currently exists.

 

 

 

DALL-E 2 DALL-E 2 is an AI-powered program developed by OpenAI that generates images from textual input. It is a more advanced version of its predecessor, DALL-E, capable of creating more complex and detailed images with higher resolution (up to 512 x 512 pixels). It uses a generative model trained on a large dataset of images and can be controlled with various parameters such as pose, color, shape, and texture.

 

Clip Guided Diffusion CLIP guided diffusion is a technique that uses the CLIP neural network to generate high-quality images by interpolating between two or more input images. It works by optimizing a set of parameters that control the diffusion process, including the diffusion step size, number of iterations, and blending factors. This approach can be used to create realistic and diverse images, and has applications in fields such as digital art, computer graphics, and generative modeling.
VQGAN-CLIP VQGAN+CLIP is a combination of two neural network architectures: VQGAN and CLIP. In one sentence: CLIP guides VQGAN towards an image that is the best match to a given text.

VQGAN is trained on a mostly canonical dataset like ImageNet or COCO. CLIP on the other hand contains 400 million parameters.

Table 2 – image generation comparison

As shown in the table, model selection, training, and prompting make a big difference in the results achieved. Humans need to tune the input and judge the results. Assumptions about model accuracy need to be tested periodically to see what has changed.

This highlights that one vital skill you’ll need in the 21st century – effectively talking to machines. And for now, that process involves writing—or, in tech vernacular, engineering—prompts.

As Charlie Warzel points out in the Atlantic, asking ChatGPT to write a five-paragraph book report about Animal Farmwill yield forgettable, even inaccurate results. But writing the introductory paragraph to the book report yourself and asking the tool to complete the essay will feed the machine valuable context. Better yet, instruct the machine, “Write a five-paragraph book report at a college level with elegant prose that draws on the history of the satirical allegorical novel Animal Farm. Reference Orwell’s ‘Why I Write’ while explaining the author’s stylistic choices in the novel.” It will yield a far more sophisticated and convincing output. The blog generated by ChatGPT in our exercise above was factual but bland. Perhaps by tuning the prompt, it might have been more interesting.

Good prompts aren’t just specific. They seem to reflect a deeper understanding of the model you are trying to manipulate. One way to think of prompt trial and error is as an attempt to glean what information the model is pulling from and how the AI organizes and indexes the information at its disposal. It’s informed guesswork as you try to understand (manipulate) the weighting and thresholds the model uses to generate the results.

For example, in DALLE 2, the prompt: Red haired girl generates this response:

While the prompt Girl with red hair generates this one:

Most people who regularly use search often have a similar experience with search queries. You type in what you are looking for and the search engine produces results. You scan the sites it recommends as answers to your query, see that they aren’t exactly what you wanted, and refine the search query. The difference is that Generative AI produces specific outputs (this is the answer) rather than providing a range of options to choose from.

Working together

AI systems and humans are complimentary; their respective strengths make up for the other’s shortcomings.

AI systems are good at digesting vast amounts of data and increasingly good at presenting their findings in creative and even inventive ways.

At the same time, such systems lack common sense, struggle to cope with novel situations, and can incorporate and automate existing bias, among other problems.

AI systems can only respond based on the data they have ingested. If it’s not in the data, it doesn’t exist.

Humans, for all their flaws, are amazing at tasks that computers still struggle with. Even toddlers readily master tasks like understanding object permanence or identifying novel obstacles.

Importantly, humans can learn by venturing out into the world and having new experiences, while AI systems are largely limited to the diet of digital information on which they are trained.

But people are also often not at their best, thanks to boredom, distraction, exhaustion, moodiness, or many other all-too-human conditions.

When working together, AI systems can do the work and humans can make sure the outputs are trustworthy. As mentioned in our report “An Operating Model for the Digital Enterprise”, while the value of artificial intelligence is now undeniable, the question has become how best to use it – and that often boils down to the extent to which workers and end users trust AI tools…and which tools and results are deemed “trustworthy”. This is especially true with Generative AI.

Implications for Enterprises

Although still in its early stages of scaling, generative AI models are beginning to demonstrate their potential for transforming a range of business functions. The possibilities are truly exciting. Some of the first applications of these models are already emerging, including the following:

  • Marketing and sales: These models are being used to create personalized marketing content, develop social media campaigns, and craft technical sales content, including text, images, and video. They can also create assistants customized for specific businesses, such as those in the retail industry.
  • Operations: Generative AI is being used to generate task lists for efficient execution of various activities and to develop chatbot support for service queries, improving the overall efficiency of operations.
  • IT/engineering: The technology is being used to write, document, and review code, streamlining the process of software development and maintenance.
  • Risk and legal: The models are being employed to answer complex questions, drawing from vast amounts of legal documentation. They can also draft and review annual reports, making the process more efficient and accurate.
  • HR: These models are being used to assist the human resources (HR) in areas like recruitment, onboarding, training, and engagement.
  • General Productivity: Generative AI can impact productivity across various functions including communications, automating repetitive tasks, knowledge sharing, and collaboration.

Getting Ready

Implementing generative AI use cases in the enterprise requires significant effort, resources, and coordination across different departments. Here are some of the key factors that organizations should consider when implementing generative AI use cases:

  • Data preparation: Generative AI models require large amounts of high-quality data to learn and produce accurate results. Organizations need to invest time and resources in collecting, cleaning, and preparing data to ensure that it is suitable for use in generative AI models.
  • Infrastructure and tools: Generative AI models require significant computational resources to train and run. Organizations need to invest in powerful hardware and software tools to support the development and deployment of generative AI models.
  • Skills and expertise: Developing and deploying generative AI models requires specialized skills and expertise in areas such as data science, machine learning, and software engineering. Organizations may need to hire or train staff with these skills or engage with external partners to support their generative AI initiatives.
  • Ethical and legal considerations: Generative AI models can have significant implications for privacy, security, and ethical considerations. Organizations need to establish clear policies and guidelines to ensure that generative AI is used in an ethical and transparent manner and complies with relevant regulations.

Table 3 shows some examples of use cases that might apply across a typical enterprise. These cases are not exhaustive and depending on the enterprise some cases may be more difficult to implement than others.

Marketing and Sales Operations IT/Engineering Risk and Legal HR Utility and Employee Optimization
Write marketing and sales copy including text, images, and videos (eg, to create social media content or technical sales content)

 

Create or improve customer support chatbots to resolve questions about products, including generating relevant cross-sell leads

 

Write code and documentation to accelerate and scale developments (eg, convert simple JavaScript expressions into Python)

 

Draft and review legal documents, including contracts and patent applications

 

Assist in creating interview questions for candidate assessment (eg, targeted to function, company philosophy, and industry)

 

Optimize communication of employees (eg, automate email responses and text translation or change tone or wording of text)

 

Create product user guides of industry-dependent offerings (eg, medicines or consumer products)

 

Identify production errors, anomalies, and defects from images to provide rationale for issues

 

Automatically generate or autocomplete data tables while providing contextual information

 

Summarize and highlight changes in large bodies of regulatory documents

 

Provide self-serve HR functions (eg, automate first-line interactions such as employee onboarding Create business presentations based on text prompts, including visualizations from text

 

Analyze customer feedback by summarizing and extracting important themes from online text and images

 

Streamline customer service by automating processes and increasing agent productivity

 

Generate synthetic data to improve training accuracy of machine learning models with limited unstructured input

 

Answer questions from large amounts of legal documents, including public and private company information

 

Develop personalized learning plans for each employee based on their skills and learning style. Synthesize a summary (eg, from text, slide decks, or online video meetings)

 

Improve sales force by, for example, flagging risks, recommending next interactions such as additional product offerings, or identifying optimal customer Identify clauses of interest, such as penalties or value owed through leveraging comparative document analysis

 

Identify and assess risks through data analysis and simulation. Develop virtual assistants that can answer common employee queries and provide guidance on company policies and procedures. Enable search and question answering on companies’ private knowledge data (eg, intranet and learning content)

 

Create or improve sales support chatbots to help potential clients under-stand, including technical product understanding, and choose products

 

Monitor compliance with regulatory requirements and detect any potential violations. Automated accounting by sorting and extracting documents using automated email openers, high-speed scanners, machine learning, and intelligent document recognition

 

Table 3 – Example Use Cases

When considering each potential use case, make sure the necessary data, infrastructure, skills, and policies are available to initiate and evolve the AI capabilities as intended.

Caveat Emptor

The awe-inspiring results of generative AI might make it seem like a ready-set-go technology, but that’s not the case. The nascent state of generative AI requires enterprises to proceed with an abundance of caution. Technologists are still working out the kinks, and plenty of practical and ethical issues remain open. Here are just a few:

  • Like humans, generative AI can be wrong. ChatGPT, for example, confidently generates inaccurate information in response to a user question and has no built-in mechanism to signal this to the user or challenge the result. For example, in our capability testing above, when the tool was asked to create the real estate listing, it generated several incorrect facts for the property, such as listing the wrong number of bedrooms and bathrooms, and a deck instead of a patio. The only way I knew this was wrong was because I lived there.

It’s not just a history problem. As the world around us constantly changes, AI systems need to be constantly retrained using new data. Without this crucial step, AI systems will produce answers that are factually incorrect or do not consider new information that’s emerged since they were trained.

  • Filters are not yet effective enough to catch inappropriate content. Users of an image-generating application that can create avatars from photos have received avatar options from the system that portrayed them nude, even though they had input appropriate (clothed) photos of themselves. Furthermore, there will be “bad actors” that will look for ways of generating inappropriate, misleading or incorrect results.
  • Systemic biases still need to be addressed. These systems draw from massive amounts of data that, just like the internet itself, might include unwanted biases.
  • Individual company norms and values aren’t reflected. Companies will need to adapt the technology to incorporate their culture and values, an exercise that requires technical expertise and computing power beyond what some companies may have ready access to.
  • Intellectual-property questions are up for debate. When a generative AI model brings forward a new product design or idea based on a user prompt, who can lay claim to it? What happens when it plagiarizes a source based on its training data? This is especially true when enterprise proprietary information augments public data.
  • Bad actors can use the power of generative AI. As the ability of generative AI becomes indistinguishable from humans, the spread of fake accounts and content will become less detectable and therefore more harmful.

Conclusion

The latest version of ChatGPT is a testament to the constant advancements being made in language model technology and is potentially game changing for individuals and enterprises. It is important to note that language models still have their limitations, are relatively immature and should not be fully relied on for important matters.

GPT4 or generative pre trained transformer four is upcoming version of OpenAI’s language model. This model is reported to process text at an unprecedented speed and its capabilities are truly game changing. OpenAI has said that the focus of GPT4 goes beyond speed; it includes improvements in accuracy and alignment as well. The highly anticipated release of GPT4 is on the horizon and while rumors and speculation surround its capabilities, one thing is certain. The partnership between Microsoft and Open AI is a game-changer.

The partnership between Microsoft and OpenAI is an inflection point for AI. With Microsoft’s vast resources, expertise in technology, and embedded base combined with OpenAI’s cutting edge AI research, we can expect to see incredible advancements in the field of generative AI in the coming years. Think about using ChatGPT to assist you in composing an outlook e-mail or a word document; the possibilities are infinite thanks to the combined knowledge and resources of these two companies. This partnership will open the door for fresh and innovative uses of AI across several sectors transforming how we live and work so pay attention and prepare for AI’s future.

The possibilities are numerous ranging from enhancing natural language processing to creating more complex machine learning models. So, while we can’t say for sure what GPT4 will bring, one thing is certain OpenAI is pushing the boundaries of AI technology and we can’t wait to see what it can do. As always, it’s important to use this powerful technology responsibly and to ensure that it is used ethically.

With Bard (and its underlying model LaMDA), Google is ready to disrupt the impact of ChatGPT with its own version, giving professionals and consumers more options to get answers. From creativity to purpose, brands can leverage all that data to enhance their presence (and profits). However, rushing to market has its perils. On 8 January 2023, Google parent company lost $100 billion in market value after Bard shared inaccurate information in a promotional video and a company event failed to dazzle, feeding worries that the Google parent is losing ground to rival Microsoft Corp. However, the story is not over as Google has a history of overcoming challenges to remain dominant in its markets.

There are other platforms out there.

Not all problems can be addressed by these large AI implementations. Enterprise specific data, norms, and values may need to go beyond these generic models into something more specific; and not everything is solved via text. Other platforms are out there for developing and training your own model and generating outputs beyond text, neither of which are currently available with GPT or LaMDA.

Hugging Face is an open-source and platform provider of machine learning technologies. Hugging Face was launched in 2016 and is headquartered in New York City. Hugging Face allows users to build, train, and deploy custom models based on user contributed open-source model references.

Image generation, as shown in table 2, has many service providers that can generate custom artwork and images for use in marketing and content illustration.

Artificial intelligence (AI) is making it easier than ever to generate video. Video generators like Hour One, Pictory, and Synthesis take text scripts and images and turn them into near production ready pieces with no experience in video editing or design needed. Here is an example generated using Hour One from a portion of the executive summary of this report.

https://vimeo.com/807043063

Understanding the fast-moving landscape and experimenting with the various capabilities are critical now; the key word is experimenting.

Recommendations

Incorporating generative AI technology into a company’s operations is a decision that requires careful consideration. To maximize the technology’s potential and stay ahead of its rapid evolution, it is crucial to identify the areas of the business where it can have the most immediate effect and establish an internal / external monitoring mechanism. A prudent first step is to establish a cross-functional (think governance) team that includes data science practitioners, legal experts, and functional business leaders, to consider fundamental questions such as the potential impact of the technology on the industry or value chain, policies, and posture, and criteria for selecting use cases.

To build an effective ecosystem of partners, communities, and platforms, the team must think beyond the limitations of the generative AI models and establish legal and community standards to maintain stakeholders’ trust. At the same time, fostering thoughtful innovation across the organization is vital. This can be achieved by creating sandboxed environments for experimentation, many of which are readily available through cloud-based solutions and implementing guardrails for experiment and use.

Generative AI is an opportunity that also comes with risks. AI is still expected to require plenty of oversight by humans. The curation, editing, and organization will take time and resource.

As the AI discussion washes over us, both in real life and online, many experts reach the same conclusion: AI is meant to assist professionals and help them do their jobs more effectively, not to replace them entirely. That’s yet to be proven.

The consensus among many experts is that several professions will be totally automated in the next five to 10 years. A group of senior-level tech executives who comprise the Forbes Technology Council named 15: insurance underwriting, warehouse and manufacturing jobs, customer service, research (yikes) and data entry, long haul trucking and a somewhat disconcertingly broad category titled “Any Tasks That Can Be Learned.”

Ultimately, it’s hard to predict exactly how AI will be used in the real world. However, the potential to reduce waste and increase the quality of content is there. ChatGPT has demonstrated its capabilities are undeniably valuable, whether it’s for writing code or breaking down complex ideas into digestible explanations.

Whatever we think the potential is for Generative AI, it will surely change as the capabilities evolve and are integrated into our work and personal lives.

About TechVision

World-class research requires world-class consulting analysts, and our team is just that. Gaining value from research also means having access to research. All TechVision Research licenses are enterprise licenses; this means everyone that needs access to content can have access to content. We know major technology initiatives involve many different skillsets across an organization and limiting content to a few can compromise the effectiveness of the team and the success of the initiative. Our research leverages our team’s in-depth knowledge as well as their real-world consulting experience. We combine great analyst skills with real world client experiences to provide a deep and balanced perspective.

TechVision Consulting builds off our research with specific projects to help organizations better understand, architect, select, build, and deploy infrastructure technologies. Our well-rounded experience and strong analytical skills help us separate the “hype” from the reality. This provides organizations with a deeper understanding of the full scope of vendor capabilities, product life cycles, and a basis for making more informed decisions. We also support vendors in areas such as product and strategy reviews and assessments, requirement analysis, target market assessment, technology trend analysis, go-to-market plan assessment, and gap analysis.

TechVision Updates will provide regular updates on the latest developments with respect to the issues addressed in this report.

About the Author

Gary Zimmerman is an experienced executive known for helping companies deliver new offers and expand markets. Accomplishments include launching four companies, 20+ products, building high-performance organizations, and generating millions in sales.

His experience at Neustar, Respect Network, and Sovrin allows him to provide a broad perspective on a variety of subjects including self-sovereign identity, blockchain, enterprise data management, and the data brokerage industry. His experience both enterprise and startup product development give him a unique perspective on innovation.

[1] Full disclosure, this bolded section about Generative AI was generated by ChatGPT, not the author.

[2] It’s actually a 3-bedroom 2-bath home.

[3] It has a patio not a deck.

[4] As impressive as it is, ChatGPT is still narrow AI, trained to ingest and interpret input text and generate the model’s best guess at output text.

[5] Due to the influence of InstructGPT

Tags:

We can help

If you want to find out more detail, we're happy to help. Just give us your business email so that we can start a conversation.

Thanks, we'll be in touch!

Stay in the know!

Keep informed of new speakers, topics, and activities as they are added. By registering now you are not making a firm commitment to attend.

Congrats! We'll be sending you updates on the progress of the conference.