The AI revolution: Large models change the AI engineering paradigm
Shang Can Technology
In July, 2023
abstract
The emergence of artificial intelligence large models has changed the training mode of artificial intelligence, improved the starting point of training, broken through the shortcomings and bottleneck of the original small models, improved the level of intelligence, and brought a new paradigm of artificial intelligence engineering. Enterprises should realize the significance of this revolutionary breakthrough, explore the application of large models to solve enterprise problems, build the enterprise Foundation model, and then build the enterprise intelligent machine matrix, to lay the foundation for the realization of dual-intelligent collaborative business.
The main discovery
Large models are not only language models, but also include multi-modal large models, which can process video, audio, images, and structured data
Large model is a revolutionary breakthrough in the field of artificial intelligence, and its intelligence emergence and generalization ability make artificial intelligence have a high intelligence
Large models can be used not only for content generation, but also for data analysis, modeling and algorithm extraction
The use of large models for data analysis and modeling will completely change the existing data analysis mode and bring a new paradigm of artificial intelligence engineering
The application of the large model can realize the idea of the foundation model, and a variety of small models and expert models can be trained on top of the foundation model
Recommendations
Enterprises should realize the revolutionary significance of the large model and take the foundation model created by the large model as the "base" for the business of dual intelligent collaboration
Explore the value of large models to the business, and find application scenarios in the difficult problems faced by enterprises
When selecting a large model, enterprises should comprehensively consider the applicability, safety and input-output ratio of a large model
To improve the ability to control large models, management ability is an indispensable ability, and technical ability is a plus
foundation model is an idea in the field of artificial intelligence. It is hoped that one model can be the foundation of other models, so that other models can be trained on top of the foundation model. Such a scenario is not fully realized before the large model is produced. It is hard to say that the large model is the only option for the foundation model, but there is no doubt that the large model can be the foundation model. As we have analyzed, a large model refers to the large number of model parameters (feature values of connections between neurons in the neural network). The generally agreed standard is more than 3 billion. At present, many Chinese enterprises say that the number of parameters is more than 1 billion, which can be regarded as a large model. In fact, the model is a relative concept. When the output accuracy of the model is higher than the random value, it indicates that the large model has appeared the emergent ability. The number of parameters at this point is the critical point. When the accuracy reaches the reliability requirements, the large model is available, and the number of parameters is large enough for the required task.
Large models were developed because of the need of AIGC, because machines have the ability to master a lot of knowledge through learning. When the amount of knowledge learned is not large enough, the model does not have enough intelligence to produce meaningful content, so a lot of training materials are needed. With the increase of training materials, the computational speed is faster and faster, and the model becomes more and more large, eventually leading to the generation of large models. Because large models are born to realize AIGC, and because most of the recent large models are mainly used for AIGC tasks, people mistakenly think that large models can only be used for content generation, while ignoring another more important value area, namely the value of the foundation model. The value of this field makes it possible to completely transform the paradigm of artificial intelligence engineering, from "training the model from a kid" to " from a college students ".
Augmented data analytics, a meaningful but less successful attempt at artificial intelligence
Since the concept of big data began to become popular, enterprises have established big data management system and big data analysis system, hoping to find the correlation and law between data through historical data and current data, and use these findings for the analysis of the current situation and the prediction of the future. However, before the Augmented data analysis technology, the design of the data analysis model mainly depends on the model designer's understanding of the business and the manual design. Manual design mode requires designers to have enough understanding of the business, so the designer's business ability and design ability will become the limitation of the data analysis model. On the one hand, the efficiency of manual design is very low, so it is impossible to complete the models required for the business in a relatively short time, and the total amount of models that can be completed manually is very limited. On the other hand, because the designer's understanding of the business is insufficient, the structure of the design cannot fully reflect the law of the business, forming a design bias. These things improved after machine learning was introduced into big data analytics. The automatic discovery of data (Pattern) through machine learning can greatly improve the efficiency of model generation, and break through the cognitive limitations of designers, so that the quantity and quality of analytical models can be improved to a greater extent. If the manual design analysis model is used, a designer can design a model in a year is also very limited. Using Augmented data analysis technology produces thousands of analytical models within minutes. Of course, it will take some time to verify the validity of these models, but the efficiency has been greatly improved.
According to the survey, a considerable number of enterprises have adopted augmenteddata analysis technology. However, the overall application effect is not ideal. There are two main reasons. One is that the technology of augmented data analysis is not well mastered and can not effectively play its effectiveness. Second, the augmented data analysis itself also has defects, because the machine can only find correlation, not causality. At the same time, the focus of machine learning is to discover data analysis models, rather than improving the capabilities of machine learning models themselves. In some application scenarios, such as the application of various intelligent brain, although improving the ability of the intelligent brain is the key, after long-term training, it is found that when the ability of the intelligent brain is improved to a certain extent, it seems to reach the ceiling of the ability, no matter how the training can not be improved. For example, the driverless system cannot overcome the difficulties of complex traffic environment, and the error rate of the language translation system kept at about 20% cannot be further reduced. In the medical field, the recognition of medical images has also hit the ceiling, always not at the level of human experts. In the f