Sat.Sep 09, 2023 - Fri.Sep 15, 2023

article thumbnail

The Role of DevOps and CI/CD in Data Engineering

Confessions of a Data Guy

In the vast world of data, it’s not just about gathering and analyzing information anymore; it’s also about ensuring that data pipelines, processes, and platforms run seamlessly and efficiently. Nothing screams “why are flying by night,” than coming into a Data Team only to find no tests, no docs, no deployments, no Docker, no nothing. […] The post The Role of DevOps and CI/CD in Data Engineering appeared first on Confessions of a Data Guy.

article thumbnail

GPT and LLMs from a Data Engineering Perspective

Jesse Anderson

There has been quite a bit of writing covering GPT and LLMs from data science and business perspectives. I haven’t seen much from the data engineering side. Let me share my perspective, having been in data and AI for a while and using LLMs before they became popular. It is interesting to see the general public having the same amount of excitement as there was a year ago in the LLM space.

Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

Data News — Week 23.37

Christophe Blefari

Facing the News ( credits ) Hello Data News readers. I'm still struggling to get back into my usual work rhythm. If you add the fact that last week I came up with fewer articles than I expected, this has led me to another blank page. Anyway, after 2 years of work, I have to accept and let go when necessary. But don't worry I don't forget you.

article thumbnail

Apache Flink best practices - Flink Forward lessons learned

Waitingforcode

I won't hide it, I'm still a fresher in the Apache Flink world and despite my past streaming experiences with Apache Spark Structured Streaming and GCP Dataflow, I need to learn. And to learn a new tool or concept, there is nothing better than watching some conference talks!

IT 130
article thumbnail

Navigating the Future: Generative AI, Application Analytics, and Data

Generative AI is upending the way product developers & end-users alike are interacting with data. Despite the potential of AI, many are left with questions about the future of product development: How will AI impact my business and contribute to its success? What can product managers and developers expect in the future with the widespread adoption of AI?

article thumbnail

How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federated Data Teams

Seattle Data Guy

Success in the data world hinges on team setup. I’ve delved into onboarding and standards in previous articles, but never into the structure of data teams. Typically, there are three configurations: Centralized, Decentralized, and Federated. Most companies I’ve seen use a mix of these. While the newest tech breakthroughs grab headlines, team organization is the… Read more The post How To Set Up Your Data Analytics Team For Success – Centralized vs Decentralized vs Federat

article thumbnail

Everything you Need to Become a SAS Certified Data Scientist

KDnuggets

With a shortage of talent and an abundance of opportunity, there’s never been a better time to launch or advance your data science career with the SAS Academy for Data Science. Read on to find out everything you need to become a SAS Certified Data Scientist.

More Trending

article thumbnail

Meta Quest 2: Defense through offense

Engineering at Meta

Meta’s Native Assurance team regularly performs manual code reviews as part of our ongoing commitment to improve the security posture of Meta’s products. In 2021, we discovered a vulnerability in the Meta Quest 2’s Android-based OS that never made it to production but helped us find new ways to improve the security of Meta Quest products. We’re sharing our journey to get arbitrary native code execution in the privileged VR Runtime service on the Meta Quest 2 by exploiting a memory corruption v

Bytes 129
article thumbnail

Measuring Technical Debt to Avoid the Boiling Frog Syndrome

Booking.com Engineering

source Software development is all about change. And, over the lifespan of our software, the goal is to implement required changes in a reasonable amount of time. Whether the changes are technical in nature, like an urgent security upgrade, or stem from a business need, such as building a new feature to make us more competitive in target markets — how fast we can change is critical.

Coding 98
article thumbnail

Getting Started with SQL in 5 Steps

KDnuggets

This comprehensive SQL tutorial covers everything from setting up your SQL environment to mastering advanced concepts like joins, subqueries, and optimizing query performance. With step-by-step examples, this guide is perfect for beginners looking to enhance their data management skills.

SQL 112
article thumbnail

Introducing MLflow 2.7 with new LLMOps capabilities

databricks

As part of MLflow 2’s support for LLMOps, we are excited to introduce the latest updates to support prompt engineering in MLflow 2.7. A.

article thumbnail

Get Better Network Graphs & Save Analysts Time

Many organizations today are unlocking the power of their data by using graph databases to feed downstream analytics, enahance visualizations, and more. Yet, when different graph nodes represent the same entity, graphs get messy. Watch this essential video with Senzing CEO Jeff Jonas on how adding entity resolution to a graph database condenses network graphs to improve analytics and save your analysts time.

article thumbnail

A Watershed Moment

ArcGIS

Updated data from the Watershed Boundary Dataset (WBD) are added to Living Atlas as new feature services.

Datasets 127
article thumbnail

Mode + ThoughtSpot recognized as Leaders in Snowflake’s 2023 Modern Marketing Data Stack awards

ThoughtSpot

We’re thrilled to announce that both ThoughtSpot and Mode ( acquired by ThoughtSpot in July 2023 ) have been recognized as Leaders in Snowflake's recent Modern Marketing Data Stack report! Given the ever-evolving landscape of modern data analytics products, organizations are looking to ThoughtSpot and Mode when seeking innovative solutions—helping them harness the power of their marketing data.

article thumbnail

Working with Big Data: Tools and Techniques

KDnuggets

Where do you start in a field as vast as big data? Which tools and techniques to use? We explore this and talk about the most common tools in big data.

article thumbnail

How to Build an Interactive Real-Time Chat Application with Websockets?

Workfall

Reading Time: 11 minutes What is Socket.io? Socket.io , a widely-used JavaScript library, offers a framework for facilitating real-time, two-way communication between web clients (like browsers) and servers. It uses WebSockets as the primary communication method but also offers fallback options such as long polling for environments where WebSockets may not be supported.

article thumbnail

Understanding User Needs and Satisfying Them

Speaker: Scott Sehlhorst

We know we want to create products which our customers find to be valuable. Whether we label it as customer-centric or product-led depends on how long we've been doing product management. There are three challenges we face when doing this. The obvious challenge is figuring out what our users need; the non-obvious challenges are in creating a shared understanding of those needs and in sensing if what we're doing is meeting those needs.

article thumbnail

Introducing Apache Spark™ 3.5

databricks

Today, we are happy to announce the availability of Apache Spark™ 3.5 on Databricks as part of Databricks Runtime 14.0. We extend our s.

article thumbnail

6 Tips for Setting the Price of Your Data Product

Snowflake

Building your data product is only the beginning. You’ve considered a wide variety of use cases, and settled on the one you’ll focus on. Maybe you’re going to help hospitals predict emergency room visits and optimize their staffing. Or you’re going to enable restaurants to reduce their food waste. Or maybe you just have some really unique data that you think might be of use to someone.

Food 96
article thumbnail

KDnuggets Top Posts for August 2023: Forget ChatGPT, This New AI Assistant Will Change the Way You Work

KDnuggets

Forget ChatGPT, This New AI Assistant Is Leagues Ahead and Will Change the Way You Work Forever • 7 Projects Built with Generative AI • Best Python Tools for Building Generative AI Applications Cheat Sheet • Harnessing ChatGPT for Automated Data Cleaning and Preprocessing • Data Scientists Need to Specialize to Survive the Tech Winter • 7 Steps to Mastering Data Cleaning and Preprocessing Techniques • 5 Ways You Can Use ChatGPT's Code Interpreter For Data Science • The Best Courses for AI from U

article thumbnail

A Talented Team, Innovative Technology, and The Opportunity to Grow. There Is No Place Like Cloudera

Cloudera

I started my current career path with Hortonworks in 2016, back when we still had to tell people what Hadoop was. Once I got to work with all the amazing open-source Apache tools I was hooked. I found Apache NiFi especially interesting. Soon after, I became a huge fan of Apache Kafka. Coupled with amazing technology was an amazing team that only grew and improved with the merger with Cloudera.

article thumbnail

Beyond the Basics of A/B Tests: Highly Innovative Experimentation Tactics You Need to Know

Speaker: Timothy Chan, PhD., Head of Data Science

Are you ready to move beyond the basics and take a deep dive into the cutting-edge techniques that are reshaping the landscape of experimentation? 🌐 From Sequential Testing to Multi-Armed Bandits, Switchback Experiments to Stratified Sampling, Timothy Chan, Data Science Lead, is here to unravel the mysteries of these powerful methodologies that are revolutionizing how we approach testing.

article thumbnail

Crossing Bridges: Reporting on NYC taxi data with RStudio and Databricks

databricks

As data enthusiasts, we love uncovering stories in datasets. With Posit’s RStudio Desktop and Databricks Lakehouse, you can analyze data with dplyr, create i.

article thumbnail

How Marriott Modernized Their Data Architecture with Snowflake

Snowflake

More than 50% of data leaders recently surveyed by BCG said the complexity of their data architecture is a significant pain point in their enterprise. Companies hampered by legacy data architectures are often plagued by a high total cost of ownership (TCO), an inability to govern data, and a lack of scalability as their data volumes grow. “As a result,” says BCG, “many companies find themselves at a tipping point, at risk of drowning in a deluge of data, overburdened with complexity and costs.

article thumbnail

The 5 Best AI Tools For Maximizing Productivity

KDnuggets

KDnuggets reviews a diverse set of 5 AI tools to help maximize your productivity. Have a look and see what our recommendations include.

126
126
article thumbnail

Career Stories: Learning and growing through mentorship and community

LinkedIn Engineering

Lekshmy has always been interested in a role in a company that would allow her to use her people skills and engineering background to help others. Working as a software engineer at various companies led her to hear about the company culture at LinkedIn. After some focused networking, Lekshmy landed her position at LinkedIn and has been continuing to excel ever since.

article thumbnail

How Embedded Analytics Gets You to Market Faster with a SAAS Offering

Start-ups & SMBs launching products quickly must bundle dashboards, reports, & self-service analytics into apps. Customers expect rapid value from your product (time-to-value), data security, and access to advanced capabilities. Traditional Business Intelligence (BI) tools can provide valuable data analysis capabilities, but they have a barrier to entry that can stop small and midsize businesses from capitalizing on them.

article thumbnail

Improve Lakehouse Security Monitoring using System Tables in Databricks Unity Catalog

databricks

As the lakehouse becomes increasingly mission-critical to data-forward organizations, so too grows the risk that unexpected events, outages, and security incidents may derail.

Systems 89
article thumbnail

Data Cloud Industry Day 2023: Your Event Guide

Snowflake

The first annual Data Cloud Industry Day is here! Data Cloud Industry Day 2023 is a free virtual event on September 28, 2023, dedicated to what’s possible for you and your industry in the world of data. From leading-edge innovations to seamless solutions to your toughest industry-specific challenges, Industry Day provides the insights and information you need to drive business value with your data.

Cloud 87
article thumbnail

Closed Source VS Open Source Image Annotation

KDnuggets

This blog strikes a comparison between open-source and closed-source image annotation tools and how it makes the life of AI model developers easy and convenient.

IT 113
article thumbnail

Path Representation in Python

Towards Data Science

Here’s why you should avoid representing paths as strings and use Pathlib instead Continue reading on Towards Data Science »

Python 94
article thumbnail

Peak Performance: Continuous Testing & Evaluation of LLM-Based Applications

Speaker: Aarushi Kansal, AI Leader & Author and Tony Karrer, Founder & CTO at Aggregage

Software leaders who are building applications based on Large Language Models (LLMs) often find it a challenge to achieve reliability. It’s no surprise given the non-deterministic nature of LLMs. To effectively create reliable LLM-based (often with RAG) applications, extensive testing and evaluation processes are crucial. This often ends up involving meticulous adjustments to prompts.

article thumbnail

The Power of a Trusted Data Lakehouse: Go Bust or Boom

databricks

Special thanks to our partners at Immuta, Alation, and Anomalo for their collaboration on the content and technical assets from this article. The.

Data 90
article thumbnail

Understanding Snowflake’s Shared Responsibility Model

Snowflake

The White House recently released the first National Cybersecurity Strategy , which among other things, holds the stewards of data accountable and shifts liability for insecure software products and services away from end users and toward vendors that are capable of taking actions to prevent bad outcomes. We are thrilled to announce both the availability of the Snowflake Shared Responsibility Model , alongside Snowflake’s collaboration with the Center for Internet Security (CIS) and the security

article thumbnail

5 Amazing & Free LLMs Playgrounds You Need to Try in 2023

KDnuggets

Explore the top 5 user-friendly platforms that provide free access to large language models, enabling you to experience the latest AI models firsthand.

article thumbnail

Should We Be Virtualizing Our Data Science Systems and—or Not?

Towards Data Science

It can be hard to navigate the pros and cons of virtualizing data science processes, but some power and performance trends cannot be… Continue reading on Towards Data Science »

article thumbnail

From Developer Experience to Product Experience: How a Shared Focus Fuels Product Success

Speaker: Anne Steiner and David Laribee

As a concept, Developer Experience (DX) has gained significant attention in the tech industry. It emphasizes engineers’ efficiency and satisfaction during the product development process. As product managers, we need to understand how a good DX can contribute not only to the well-being of our development teams but also to the broader objectives of product success and customer satisfaction.