Google DeepMind Unveils Gemini Robotics

MMS Founder
MMS Daniel Dominguez

Google DeepMind has introduced Gemini Robotics, an advanced AI model designed to enhance robotics by integrating vision, language, and action. This innovation, based on the Gemini 2.0 framework, aims to make robots smarter and more capable, particularly in real-world settings.

One of the key features of Gemini Robotics is its embodied reasoning, which allows robots to understand and react to their environment in a more human-like way. This capability is crucial for robots to adapt quickly in dynamic and unpredictable environments. Gemini Robotics enables robots to perform a wider range of tasks with greater precision and adaptability, which are significant advancements in robotic dexterity.

Google DeepMind is also developing the next generation of humanoid robots partnering with Apptronik, which have the potential to work alongside humans in various environments, including homes and offices. The concept of steerability is emphasized, referring to the responsiveness of robots to human commands and environmental changes, enhancing their versatility and ease of use.

Safety and ethics are top priorities, with measures such as collision avoidance and force limitation integrated into the AI models. The ASIMOV dataset, inspired by Isaac Asimov’s Three Laws of Robotics, aims to improve safety in robotic actions, ensuring robots operate ethically and safely around humans.

Comments from various sources reflect excitement and optimism highlighting its adaptability and generalization, calling it a step toward genuine usefulness in robotics, moving beyond mere automation.

Educator and business leader Patrick Egbunonu posted on X:

Imagine robots intuitively packing lunchboxes, handling delicate items, or assembling products efficiently—without extensive custom programming.

Others note its impressive dexterity and instruction-following, suggesting it could be a pivotal advancement. Web discussions, like those on Reddit, draw parallels to a ChatGPT moment for robotics, though some argue it needs broader consumer access to truly revolutionize the field. 

User ogMackBlack shared on Reddit:

The ChatGPT moment in robotics, to me at least, will be the moment regular people like us will be able to purchase them robots for personal use or have Gemini taking control of physical stuff autonomously at home via an app.

Google DeepMind’s work expands the capabilities of robotics technology, pushing its development forward. While experts recognize its potential to connect cognitive processing with physical action, some remain skeptical about its immediate real-world impact, especially when compared to high-profile demonstrations from competitors like Tesla’s Optimus.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Latin America Launches Latam-GPT to Improve AI Cultural Relevance

MMS Founder
MMS Daniel Dominguez

Latin America is advancing in the development of artificial intelligence with the creation of Latam-GPT, a language model designed to better represent the history, culture, and linguistic diversity of the region. Announced at the Paris AI Action Summit, the project is led by Chile’s Ministry of Science, Technology, Knowledge, and Innovation (CTCI) and the National Center for Artificial Intelligence (Cenia), with support from experts and institutions across Latin America.

The primary issue addressed by Latam-GPT is the misinterpretation or generalization of phrases, idioms, and cultural references common in current AI models when applied to the Latin American context. This model has been developed with 50B parameters and a database, comparable to OpenAI’s GPT-3.5. The focus on Latin American content allows for more accurate responses to historical, social, and cultural inquiries specific to the region.

The development of Latam-GPT has been funded through an investment, with contributions from the Andean Development Corporation and the University of Tarapacá. Collaborating countries include Argentina, Colombia, Ecuador, Mexico, Peru, and Uruguay, alongside support from institutions in the United States and Spain.

As an open-source initiative, it enables developers and researchers to tailor applications to regional requirements. Its potential uses span education, public policy analysis, environmental research, and economic studies. The model’s open nature fosters innovation and adaptation to local challenges. Currently in the training phase, it invites public participation by collecting Spanish texts that capture linguistic nuances from different countries. This crowdsourced approach helps refine the model’s understanding of regional dialects and cultural contexts.

Comments have highlighted interest and discussions around Latam-GPT, indicating a blend of optimism about fostering digital sovereignty and skepticism about its practical applications beyond educational and cultural research.

User Daniel Ramirez commented on X:

LatamGPT: The First AI Model Made in Latin America Developed by @cen_ia, trained with local data, and featuring 50B parameters, it aims to reduce biases and improve accuracy in the region. Will it be key to digital sovereignty?

And user quasar_buckhead shared:

Bravo Latin America, Bravo! This is exactly what I have been talking about for the last month! There is no “One AI LLM”. Every AI nation should develop their own LLMs for their own needs and disciplines! Excellent Work Latin America! Bravo!

Latam-GPT represents a significant step toward enhancing technological sovereignty and reducing reliance on external AI solutions in Latin America. By focusing on regional specificity, the model aims to foster more inclusive and equitable technological advancements, reflecting the rich diversity of Latin America’s culture and history. Initiatives like Latam-GPT and OpenEuroLLM highlight regional efforts to develop AI models tailored to local cultural and linguistic contexts. Other similar initiatives include SeaLLMs for Southeast Asia, SEA-LION for Singapore, Mistral Saba for Arabic-speaking countries, and Orange’s African Language Model for regional African languages.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Hugging Face Expands Serverless Inference Options with New Provider Integrations

MMS Founder
MMS Daniel Dominguez

Hugging Face has launched the integration of four serverless inference providers Fal, Replicate, SambaNova, and Together AI, directly into its model pages. These providers are also integrated into Hugging Face’s client SDKs for JavaScript and Python, allowing users to run inference on various models with minimal setup.

This update enables users to select their preferred inference provider, either by using their own API keys for direct access or by routing requests through Hugging Face. The integration supports different models, including DeepSeek-R1, and provides a unified interface for managing inference across providers.

Developers can access these services through the website UI, SDKs, or direct HTTP calls. The integration allows seamless switching between providers by modifying the provider name in the API call while keeping the rest of the implementation unchanged. Hugging Face also offers a routing proxy for OpenAI-compatible APIs.

Rodrigo Liang, co-founder & CEO at SambaNova, stated:

We are excited to be partnering with Hugging Face to accelerate its Inference API. Hugging Face developers now have access to much faster inference speeds on a wide range of the best open source models.

And Zeke Sikelianos, founding designer at Replicate, quoted:

Hugging Face is the de facto home of open-source model weights, and has been a key player in making AI more accessible to the world. We use Hugging Face internally at Replicate as our weights registry of choice, and we’re honored to be among the first inference providers to be featured in this launch.

Fast and accurate AI inference is essential for many applications, especially as demand for more tokens increases with test-time compute and Agentic AI. Open-source models help optimize performance on RDU, enabling developers to achieve up to 10x faster inference with improved accuracy.

Billing is handled by the inference provider if a user supplies their own API key. If requests are routed through Hugging Face, charges are applied at standard provider rates with no additional markup.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Stability AI Announces Integration of Top Text-to-Image Models with Amazon Bedrock

MMS Founder
MMS Daniel Dominguez

Stability AI has introduced three new text-to-image models to Amazon Bedrock: Stable Image Ultra, Stable Diffusion 3 Large, and Stable Image Core. These models focus on improving performance in multi-subject prompts, image quality, and typography. They are designed to generate high-quality visuals for various use cases in marketing, advertising, media, entertainment, retail, and more.

The new models offer a range of features: Stable Image Ultra delivers high-quality, photorealistic outputs, making it ideal for professional print media and large-format applications; Stable Diffusion 3 Large balances generation speed and output quality, suited for producing high-volume, high-quality digital assets such as websites, newsletters, and marketing materials; and Stable Image Core is optimized for fast, affordable image generation, perfect for rapidly iterating on concepts during the ideation phase.

These models address common challenges like rendering realistic hands and faces and offer advanced prompt understanding for spatial reasoning, composition, and style.

Stability AI Diffusion models, can be used for inference calls using the InvokeModel and InvokeModelWithResponseStream operations. These models support various input and output modalities, which can be found in the Stability AI Diffusion prompt engineering guide.

To use a Stability AI Diffusion model, the model ID is required, which can be found in the Amazon Bedrock model IDs. Some models also work with the Converse API, which can be checked in the Supported models and model features.

It is important to note that the Stability AI Diffusion models support different features and are available in specific AWS Regions, which can be found in Model support by feature and Model support by AWS Region respectively.

Stuart Clark, senior developer advocate at AWS, shared on his X profile:

Imagine creating stunning visuals with just text prompts, now you can!

User AI Master Tools commented:

AI in Creative Fields: Stability AI’s launch of a new model to Amazon BedRock and the debut of Luma AI’s Dream Machine 1.6 highlight the ongoing integration of AI into creative processes, from art to content generation.

The community’s response to Stability AI’s integration of its top three text-to-image models into Amazon Bedrock has been diverse, reflecting a mix of excitement, strategic insight, and critical perspectives. Enthusiasts are eager to see how this development will transform content creation across industries, while others are focused on the technical advantages and increased accessibility this integration provides. However, concerns about centralization, data privacy, and the impact on open-source AI remain part of the conversation.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

NVIDIA NIM Now Available on Hugging Face with Inference-as-a-Service

MMS Founder
MMS Daniel Dominguez

Hugging Face has announced the launch of an inference-as-a-service capability powered by NVIDIA NIM. This new service will provide developers easy access to NVIDIA-accelerated inference for popular AI models.

The new service allows developers to rapidly deploy leading large language models such as the Llama 3 family and Mistral AI models with optimization from NVIDIA NIM microservices running on NVIDIA DGX Cloud. This will help developers quickly prototype with open-source AI models hosted on the Hugging Face Hub and deploy them in production.

The Hugging Face inference-as-a-service on NVIDIA DGX Cloud powered by NIM microservices offers easy access to compute resources that are optimized for AI deployment. The NVIDIA DGX Cloud platform is purpose-built for generative AI and provides scalable GPU resources that support every step of AI development, from prototype to production.

To use the service, users must have access to an Enterprise Hub organization and a fine-grained token for authentication. The NVIDIA NIM Endpoints for supported Generative AI models can be found on the model page of the Hugging Face Hub.

Currently, the service only supports the chat.completions.create and models.list APIs, but Hugging Face is working on extending this while adding more models. Usage of Hugging Face Inference-as-a-Service on DGX Cloud is billed based on the compute time spent per request, using NVIDIA H100 Tensor Core GPUs.

Hugging Face is also working with NVIDIA to integrate the NVIDIA TensorRT-LLM library into Hugging Face’s Text Generation Inference (TGI) framework to improve AI inference performance and accessibility. In addition to the new Inference-as-a-Service, Hugging Face also offers Train on DGX Cloud, an AI training service.

Clem Delangue, CEO at Hugging Face, posted on his X account:

Very excited to see that Hugging Face is becoming the gateway for AI compute!

And Kaggle Master Rohan Paul shared a post on X saying:

So, we can use open models with the accelerated compute platform of NVIDIA DGX Cloud for inference serving. Code is fully compatible with OpenAI API, allowing you to use the openai’s sdk for inference.

At SIGGRAPH, NVIDIA also introduced generative AI models and NIM microservices for the OpenUSD framework to accelerate developers’ abilities to build highly accurate virtual worlds for the next evolution of AI.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

AWS Introduces Amazon Q Developer in SageMaker Studio to Streamline ML Workflows

MMS Founder
MMS Daniel Dominguez

AWS announced that Amazon SageMaker Studio now includes Amazon Q Developer as a new capability. This generative AI-powered assistant is built natively into SageMaker’s JupyterLab experience and provides recommendations for the best tools for each task, step-by-step guidance, code generation, and troubleshooting assistance.

Amazon Q Developer is designed to simplify and accelerate the ML development lifecycle by allowing users to build, train, and deploy ML models without leaving SageMaker Studio to search for sample notebooks, code snippets, and instructions. It can help with translating complex ML problems into smaller tasks and searching for relevant information in documentation.

The assistant is capable of generating code for various ML tasks, such as training an XGBoost algorithm for prediction or downloading a dataset from S3 and reading it using Pandas. It can also provide guidance for debugging and fixing errors, as well as recommendations for scheduling a notebook job.

JupyterLab in SageMaker Studio can now kick off models development lifecycle with Amazon Q Developer. It allows chat capability to discover and learn how to leverage SageMaker features for use cases without having to sift through extensive documentation. The assistant can also generate code tailored to the user’s needs and provide in-line code suggestions and conversational assistance to edit, explain, and document code in JupyterLab.

Ricardo Ferreira, DevRel for AWS, shares on his X account:

Silly coding mistakes are okay when you’re learning a programming language. But not so much as you progress in your software development career. #AmazonQDeveloper can help you with this.

AWS Developer Advocate Romain Jourdan, posted on X:

The generative AI space is moving so fast that it is difficult to catch up. Amazon Q Developer is improving every week too so we wanted to make it easy for developers to know what’s new, and what to test.

Other similar tools include RapidMiner, H2O.ai, KNIME, and Alteryx. These tools offer automated machine learning, data preparation, and model deployment capabilities, and can help streamline the development process and increase productivity.

Amazon Q Developer is now available in all regions where Amazon SageMaker is generally available. It is available for all Amazon Q Developer Pro Tier users, with pricing information available on the Amazon Q Developer pricing page.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Recap of MSBuild 2024: Copilot AI Agents, Phi-3, GPT-4o on Azure AI

MMS Founder
MMS Daniel Dominguez

Microsoft recently held its annual MSBuild developer conference, where it made several significant announcements, including updates to its AI capabilities, focusing on Copilot AI Agents, Phi-3, and GPT-4o now available on Azure AI.

New features for Microsoft Copilot were announced aimed at enhancing productivity and collaboration across organizations. The updates include Team Copilot, which expands Copilot’s role from a personal assistant to a team collaborator, facilitating meetings, managing tasks, and improving group communication in tools like Microsoft Teams and Microsoft Planner.

According to professor Ethan Mollick:

Agents represent the first break away from the chatbot and copilot models for interacting with AI.

Additionally, custom agents built with Microsoft Copilot Studio can now automate business processes, reason over user actions, and learn from feedback, aiming to boost efficiency and cost savings. New Copilot extensions and connectors allow developers to tailor and integrate Copilot with specific business systems, using Copilot Studio or Teams Toolkit for Visual Studio.

Microsoft also introduced Phi-3, a family of small open models developed by Microsoft. These models support developers in building cost-efficient and responsible multimodal generative AI applications. Phi-3-mini, Phi-3-small, Phi-3-medium, and Phi-3-vision are all super-sets of previous versions, offering a range of capabilities for various applications.

As mentioned by Machine Learning Researcher Awni Hannun on X:

You can run Phi3 Small (7B) in MLX LM.

The model has a few quirks: the block sparse attention, a new nonlinearity, and an unusual ways of splitting queries / keys / values.

Useful to have a flexible framework to implement it in. And still runs quite fast on an M2 Ultra

The previously available Phi-3-mini and Phi-3-medium models can now be accessed via Azure AI’s models as a service offering. Phi-3 models, optimized for various hardware and scenarios, offer cost-effective solutions for language, reasoning, and coding tasks. Notable use cases include ITC’s AI copilot for farmers, Khan Academy’s math tutoring, and Epic’s patient history summaries.

Finally, OpenAI’s GPT-4o, a new multimodal model, is now available in Azure AI Studio. This model, which is an extension of GPT-4, allows for a richer user experience by enabling inputs and outputs that span across text, images, and more. Azure OpenAI Service customers can explore GPT-4o’s capabilities in a preview playground in Azure OpenAI Studio, available in two US regions. GPT-4o is engineered for speed and efficiency, offering advanced handling of complex queries with minimal resources, translating to cost savings and improved performance.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

Databrix Announces DBRX, an Open Source General Purpose LLM

MMS Founder
MMS Daniel Dominguez

Databricks launched DBRX, a new open-source large language model (LLM) that aims to redefine the standards of open models and outperform well-known competitors on industry benchmarks.

With 132 billion parameters, DBRX has demonstrated in its own run of industry benchmarks that this model outperforms popular open-source LLMs such as LLaMA 2 70B, Mixtral, and Grok-1 across various language understanding, programming, and math tasks. The new model even competes favorably against Anthropic’s closed-source model Claude on specific benchmarks.

The AI Community expressed excitement about the release of DBRX, with Clem Delangue, CEO at Hugging Face posting on X:

Not a surprise but DBRX is already #1 trending on HF!

DBRX’s performance is attributed to its more efficient mixture-of-experts architecture, making it up to 2x faster at inference than LLaMA 2 70B despite having fewer active parameters. Databricks claims that training the model was also approximately 2x more compute-efficient than dense alternatives.

The model was pretrained on 12 trillion tokens of curated text and code data, leveraging advanced technologies like rotary position encodings and curriculum learning during pretraining. Developers can interact with DBRX via APIs or utilize Databricks’ tools to fine-tune the model on their proprietary data. Integration into Databricks’ AI products is already underway. DBRX is available on GitHub and Hugging Face.

Databricks anticipates that developers will adopt the model as a foundation for their own LLMs, potentially enhancing customer chatbots or internal question-answering systems. This approach also provides insight into how DBRX was constructed using Databricks’ proprietary tools.

To create the dataset utilized in developing DBRX, Databricks utilized Apache Spark and Databricks Notebooks for data processing, Unity Catalog for data management and governance, and MLflow for experiment tracking.

DBRX sets a new standard for open-source AI models, offering customizable and transparent generative AI solutions for enterprises. A recent survey from Andreessen Horowitz, indicates a growing interest among AI leaders to increase open-source adoption as fine-tuned models approach closed-source performance levels. Databricks expects DBRX to accelerate the shift from closed to open-source solutions.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.

xAI Opens Grok as an Open-Source Model

MMS Founder
MMS Daniel Dominguez

Elon Musk announced that xAI would make its AI chatbot Grok open source, and now the release is accessible on GitHub and Hugging Face. This move enables researchers and developers to expand upon the model, influencing how xAI evolves Grok in the face of competition from tech giants like OpenAI, Meta, Google, Microsoft, and others. This milestone marks a significant turn in the field of AI, allowing other developers and experts in the field to access Grok’s code and related data for analysis and development.

The release of Grok as open source is a bold step that will open up new opportunities in AI research and development. Previously, industry-leading models like Mistral AI’s Mixtral and Meta’s Llama 2 dominated the AI research landscape. However, Grok stands out for its colossal size, boasting an impressive set of 314 billion parameters, nearly four times larger than its closest competitor, Llama 2.

This massive size suggests promising possibilities in terms of model accuracy and interaction capability. Grok’s weights, which are essential for its operation, are available for download, enabling developers to experiment with its structure and behavior.

@Gradio shares an X Post on all things essential about xAI’s Grok-1 release:

Now that Grok1 is open-sourced, its time we learn more about the model. All things essential about XAI’s Grok-1 release: 314B params – 8x33B MoE – 25% weights active on a token Base Model (Little) Better than Llama2 & GPT3.5 Apache2 Built in JAX & RUST 8-bit weights.

Elon Musk’s decision to take an open-source approach with Grok responds to the growing demand for transparency and collaboration in the field of AI. By sharing the code and data, Musk not only fosters innovation but also promotes accountability and public evaluation of the model.

Seeking an alternative to OpenAI and Google, Musk launched xAI with the aim of developing what he described as an AI focused on maximizing truth-seeking capabilities.

Open-source software, like Grok, offers a range of benefits for both developers and the community at large. Firstly, it allows for greater transparency and auditing, contributing to the trust and reliability of the software. Additionally, it fosters collaboration and knowledge sharing among developers worldwide, which can accelerate the pace of innovation.

About the Author

Subscribe for MMS Newsletter

By signing up, you will receive updates about our latest information.

  • This field is for validation purposes and should be left unchanged.