multimodal AI
Articles tagged multimodal AI on mistr.AI.
- Seven AI breakthroughs in March 2026 that are pushing the boundaries of artificial intelligence — March 2026 brought a fundamental shift in artificial intelligence. AI systems are no longer just text generators — they are transforming into autonomous agents that plan, decide, and act. The cost of running models is falling, robots are learning to operate in the real world, and language models are beginning to understand code and security at the level of experienced developers. Here is an overview of the seven most important breakthroughs of the month.
- Google's Nano Banana 2 Combines Local AI with Cloud Power and Targets Creators and Small Businesses — Google has presented the second generation of its compact AI device Nano Banana 2. Built on the Gemini Flash model, it promises professional generative AI features completely free of charge and aims to directly compete with paid alternatives. How does it all work, what can it do, and where are the open questions?
- Perplexity Creates New AI Agents Division with Team from Visual Electric Startup — Search startup Perplexity is expanding into the AI agents space. It is acquiring the team from Visual Electric, a Sequoia-backed AI image generation tool. Founders with experience from Facebook, Apple, and Microsoft will strengthen the newly created Agent Experiences group. What does this mean for the future of AI assistants, and why will Visual Electric shut down within three months?
- 10+1 AI Trends for 2026: A Practical Guide for Czech Companies — 2025 was the year of ChatGPT experiments and pilot projects. 2026? That will be the year of hard truths. Gartner warns that 60% of organizations will fail to extract value from AI due to a chaotic approach. While some are still approving tools, others are already deploying autonomous AI agents that independently handle complex tasks. I have prepared an overview of 10+1 key trends that will separate successful companies from those that fall behind. Some you already know, others may surprise you, but you should take all of them seriously.
- South Korea Invests Billions in Homegrown Artificial Intelligence — South Korea has allocated nearly $400 million for the development of its own large language models. Five selected companies have been tasked with creating AI that will compete with OpenAI or Google. The government does not plan to fund all of them for the same duration. Every six months, results will be evaluated and the most successful will move forward. In the end, only two winners will remain. What is behind this strategy and how are the local players approaching it?
- New features in Firefly: More realistic AI videos and sounds in a few clicks — Adobe Firefly is introducing a major expansion of its AI video capabilities — it brings more realistic video generation, integrates advanced models like Veo 3, and now handles the creation of sound effects directly from voice input. What does the new feature offer and what does it mean for creatives?
- With Google's Gemini, Your Car Will Be Smarter and Safer — Google is introducing Gemini AI for Android Auto and vehicles with Google Built-In. The new digital assistant promises safer and more comfortable driving thanks to natural conversation, automatic translations, and integration with popular apps. What will this innovation bring to drivers and how can it change the travel experience?
- Google Unveiled AI That Thinks Deeper, Shops Smarter, and Creates Videos with Dialogue — Google at the I/O 2025 conference showed how far artificial intelligence can go. New Gemini 2.5, AI Mode in search, video generation with dialogue, and virtual clothes fitting — all of this promises fundamental changes in how we work, search for information, and shop. What exactly can this new generation of AI do?
- Google's New AI Mode in Search: Advanced Answers, Personalization and Ads — Google is launching AI Mode in its search, bringing not only deeper and more personalized answers, but also a new form of advertising directly in AI-generated results. What changes does this bring for users and businesses, and what can we expect from this innovation?
- Why Artificial Intelligence Still Doesn't Understand the Word "No" in Visual-Language Tasks — Artificial intelligence systems that combine images and language can now recognize objects and generate descriptions. But add a single word — "no" — and their performance drops to chance. Why do VLM models struggle so much with negation, how does it affect practice, and what can be done about it? The new NegBench benchmark has the answers.
- Google AI Studio: When You Want to Get Started with AI Quickly and Without Unnecessary Barriers — Google AI Studio is a place where anyone with an idea can play with artificial intelligence — whether you are a developer, entrepreneur, or just a curious enthusiast. What is it like when you want to try building your own AI application in the cloud?
- With Google's AI Agents, You No Longer Have to Search — Summarized Web Information Comes to You
- ChatGPT Plus Gets a Massive Upgrade: New Limits Bring Double the Messages with the Latest Models — OpenAI is significantly increasing limits for ChatGPT Plus, Team, and Enterprise users. You can now use the latest ChatGPT o3 and o4-mini models much more frequently — with full access to all tools, including web search, file analysis, and image generation. What exactly is changing and why is this an important step for everyone who wants to get the most out of ChatGPT?