Sat.Dec 24, 2022 - Fri.Dec 30, 2022

article thumbnail

I asked ChatGPT to write a blog post about Data Engineering. Here it is.

Confessions of a Data Guy

Data engineering is a vital field within the realm of data science that focuses on the practical aspects of collecting, storing, and processing large amounts of data. It involves designing and building the infrastructure to store and process data, as well as developing the tools and systems to extract valuable insights and knowledge from that […] The post I asked ChatGPT to write a blog post about Data Engineering.

article thumbnail

Should We Get Rid Of ETLs?

Seattle Data Guy

AWS has jumped on the bandwagon of removing the need for ETLs. Snowflake announced this both with their hybrid tables and their partnership with Salesforce. Now, I do take a little issue with the naming “Zero ETLs”. Because at the very surface the functionality described is often closer to a zero integration future, which probably… Read more The post Should We Get Rid Of ETLs?

AWS 130
Insiders

Sign Up for our Newsletter

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

article thumbnail

Data Science Minimum: 10 Essential Skills You Need to Know to Start Doing Data Science

KDnuggets

Data science is ever-evolving, so mastering its foundational technical and soft skills will help you be successful in a career as a Data Scientist, as well as pursue advance concepts, such as deep learning and artificial intelligence.

article thumbnail

Data Catalog - A Broken Promise

Data Engineering Weekly

Data catalogs are the most expensive data integration systems you never intended to build. Data Catalog as a passive web portal to display metadata requires significant rethinking to adopt modern data workflow, not just adding “modern” in its prefix. I know that is an expensive statement to make😊 To be fair, I’m a big fan of data catalogs, or metadata management , to be precise.

article thumbnail

Get Better Network Graphs & Save Analysts Time

Many organizations today are unlocking the power of their data by using graph databases to feed downstream analytics, enahance visualizations, and more. Yet, when different graph nodes represent the same entity, graphs get messy. Watch this essential video with Senzing CEO Jeff Jonas on how adding entity resolution to a graph database condenses network graphs to improve analytics and save your analysts time.

article thumbnail

What is Apache Arrow? Asking for a friend.

Confessions of a Data Guy

We’ve all been in that spot, especially in tech. You wanted to fit in, be cool, and look smart, so you didn’t ask any questions. And now it’s too late. You’re stuck. Now you simply can’t ask … you’re too afraid. I get it. Apache Arrow is probably one of those things. It keeps popping […] The post What is Apache Arrow?

IT 130
article thumbnail

Building a Future in Banking and Capital Markets

The Modern Data Company

Banking and Capital Markets are undergoing a period of transformation. The global economic outlook is somewhat fragile, but banks are in an excellent position to survive and thrive as long as they have the right tools in place. According to Deloitte’s report 2023 Banking and Capital Markets Outlook , banks must find ways to adapt to global disruption and understand the changing needs of consumers to find success.

Banking 52

More Trending

article thumbnail

Top 5 Data Engineering Deep Dives in 2022

Monte Carlo

No one wants to read marketing fluff, especially not data engineers. These builders and architects are prone to scoff at any article detailing concepts at a “high-level.” Everyone understands that data lineage and data pipeline monitoring are important, but the real question is, “how do you build it?” Caveat emptor, the following articles are for the technically inclined and definitely not for the faint of heart.

article thumbnail

The Terms and Conditions of a Data Contract are Data Tests

DataKitchen

The Terms and Conditions of a Data Contract are Automated Production Data Tests. A data contract is a formal agreement between two parties that defines the structure and format of data that will be exchanged between them. Data contracts are a new idea for data and analytic team development to ensure that data is transmitted accurately and consistently between different systems or teams.

article thumbnail

Snowflake: SSE File Encryption using AWS KMS

Cloudyard

Read Time: 3 Minute, 2 Second SSE File Encryption: During this post we will discuss an ERROR while executing the COPY command. Recently we got an issue while loading data from S3 bucket to Snowflake. According to the scenario, there were two files present in the bucket but surprisingly COPY command was failing to process one File. The command was reporting Access denied error for particular file.

AWS 52
article thumbnail

Top 38 Python Libraries for Data Science, Data Visualization & Machine Learning

KDnuggets

This article compiles the 38 top Python libraries for data science, data visualization & machine learning, as best determined by KDnuggets staff.

article thumbnail

Understanding User Needs and Satisfying Them

Speaker: Scott Sehlhorst

We know we want to create products which our customers find to be valuable. Whether we label it as customer-centric or product-led depends on how long we've been doing product management. There are three challenges we face when doing this. The obvious challenge is figuring out what our users need; the non-obvious challenges are in creating a shared understanding of those needs and in sensing if what we're doing is meeting those needs.

article thumbnail

Holiday Downtime, Without Data Downtime

The Modern Data Company

Data center downtime can be costly. Gartner estimates that downtime can cost $5,600 per minute, extrapolating to well over $300K per hour. When your organization’s digital service is interrupted, it can impact employee productivity, company reputation, and customer loyalty. It can also result in the loss of business, data, and revenue. With the heart of the holiday season happening, we have tips on how to enjoy holiday downtime while avoiding the high costs of data center downtime.

article thumbnail

How Data Products Are Changing Market Economics

Acceldata

From a build perspective, data products ultimately translate into products that utilize data to improve services and overall functionality. And if we were to go by this definition, it becomes clear that no product in the world can truly survive unless they are a “data product”.

article thumbnail

How to Solve 4 Elasticsearch Performance Challenges at Scale

Rockset

Scaling Elasticsearch Elasticsearch is a NoSQL search and analytics engine that is easy to get started using for log analytics, text search, real-time analytics and more. That said, under the hood Elasticsearch is a complex, distributed system with many levers to pull to achieve optimal performance. In this blog, we walk through solutions to common Elasticsearch performance challenges at scale including slow indexing, search speed, shard and index sizing, and multi-tenancy.

article thumbnail

Key Data Science, Machine Learning, AI and Analytics Developments of 2022

KDnuggets

It's the end of the year, and so it's time for KDnuggets to assemble a team of experts and get to the bottom of what the most important data science, machine learning, AI and analytics developments of 2022 were.

article thumbnail

Beyond the Basics of A/B Tests: Highly Innovative Experimentation Tactics You Need to Know

Speaker: Timothy Chan, PhD., Head of Data Science

Are you ready to move beyond the basics and take a deep dive into the cutting-edge techniques that are reshaping the landscape of experimentation? 🌐 From Sequential Testing to Multi-Armed Bandits, Switchback Experiments to Stratified Sampling, Timothy Chan, Data Science Lead, is here to unravel the mysteries of these powerful methodologies that are revolutionizing how we approach testing.

article thumbnail

Our Top 5 Most Popular Data Engineering Articles In 2022

Monte Carlo

The Pareto Principle , which holds 80% of the results will derive from 20% of the cases, is tough to escape. It definitely holds true for our Data Downtime blog with these five articles driving a majority of our traffic in 2022. There are a few characteristics that separate these articles from the chaff, namely: They were among the first to describe or even define a nascent concept.

article thumbnail

Best of 2022: Round Up

Precisely

As 2022 wraps up, we would like to recap our top posts of the year in Data Integrity, Data Integration, Data Quality, Data Governance, Location Intelligence, SAP Automation, and how data affects specific industries. Let’s take a look! Best of Data Integrity Data integrity empowers your businesses to make fast, confident decisions based on trusted data that has maximum accuracy, consistency, and context.

article thumbnail

Data Engineering Weekly in Year 2022

Data Engineering Weekly

The holidays bring joy and memories. It is always a joyful memory for me every week when I pen down (or key down 🤷🏽‍♂️) every edition of Data Engineering Weekly. I want to take a holiday break for this week's edition, and instead, I want to reflect on our journey in 2022. A Growth To Remember 2022 has been a remarkable year in terms of subscriber growth.

article thumbnail

Data-Driven Holiday Cheer: How Santa is Using Analytics to Make the Season Bright

KDnuggets

Want to know how Santa might use data science to make his job easier? So did we, so we asked ChatGPT. Read on to find out what it said.

article thumbnail

Peak Performance: Continuous Testing & Evaluation of LLM-Based Applications

Speaker: Aarushi Kansal, AI Leader & Author and Tony Karrer, Founder & CTO at Aggregage

Software leaders who are building applications based on Large Language Models (LLMs) often find it a challenge to achieve reliability. It’s no surprise given the non-deterministic nature of LLMs. To effectively create reliable LLM-based (often with RAG) applications, extensive testing and evaluation processes are crucial. This often ends up involving meticulous adjustments to prompts.

article thumbnail

Barr Moses: My Top 5 Articles of 2022

Monte Carlo

I don’t see myself as a writer or blogger. In fact, the first blog post I published on Medium sat as a draft for months. ( Data downtime , anyone?) Prior to launching Monte Carlo, I interviewed hundreds of data leaders. I gained so much insight into their hopes, dreams, and fears that the impulse to share finally exceeded the anxiety of publishing. And there was no turning back.

article thumbnail

Simple And Scalable Encryption Of Data In Use For Analytics And Machine Learning With Opaque Systems

Data Engineering Podcast

Summary Encryption and security are critical elements in data analytics and machine learning applications. We have well developed protocols and practices around data that is at rest and in motion, but security around data in use is still severely lacking. Recognizing this shortcoming and the capabilities that could be unlocked by a robust solution Rishabh Poddar helped to create Opaque Systems as an outgrowth of his PhD studies.

article thumbnail

Data News — must-read 2022 articles

Christophe Blefari

kitsch moment, from me to you ( credits ) Hey you, this is the last article of the year and it's gonna be about the articles and trends that made 2022 according to me. You'll see articles that I've already share during the year. 💡 You can also read the 2021's must-read that I've done one year and half ago or how to learn data engineering that contains key articles to understand the field.

article thumbnail

The Zen of Python

KDnuggets

Python is one of the programming languages that are very versatile and relatively easy to learn. Hence it is the choice of many new programmers, regardless of what area of tech they are interested in. It is particularly popular in all data science branches.

Python 108
article thumbnail

Entity Resolution Checklist: What to Consider When Evaluating Options

Are you trying to decide which entity resolution capabilities you need? It can be confusing to determine which features are most important for your project. And sometimes key features are overlooked. Get the Entity Resolution Evaluation Checklist to make sure you’ve thought of everything to make your project a success! The list was created by Senzing’s team of leading entity resolution experts, based on their real-world experience.

article thumbnail

Looking to the Future – How a Data Operating System Breathes Life Into Healthcare

The Modern Data Company

Looking to the Future – How a Data Operating System Breathes Life Into Healthcare Download (PDF) The post Looking to the Future – How a Data Operating System Breathes Life Into Healthcare appeared first on TheModernDataCompany.

article thumbnail

Using Product Driven Development To Improve The Productivity And Effectiveness Of Your Data Teams

Data Engineering Podcast

Summary With all of the messaging about treating data as a product it is becoming difficult to know what that even means. Vishal Singh is the head of products at Starburst which means that he has to spend all of his time thinking and talking about the details of product thinking and its application to data. In this episode he shares his thoughts on the strategic and tactical elements of moving your work as a data professional from being task-oriented to being product-oriented and the long term i

Data Lake 130
article thumbnail

How to Execute Linux Commands in Python?

Workfall

Reading Time: 8 minutes As of this writing, Linux has a global desktop market share of 2.77% ( A Report by Statcounter ), but it powers over 90% of all cloud infrastructure and hosting services. It is critical to be familiar with common Linux commands for this reason alone. According to a 2022 StackOverflow survey , Linux-based operating systems are more popular than macOS, demonstrating the appeal of using open-source software by professional developers, with an impressive 39.89% market share.

Python 52
article thumbnail

5 Tasks To Automate With Python

KDnuggets

Here are 5 tasks you can automate with Python, and how to do it.

Python 160
article thumbnail

From Developer Experience to Product Experience: How a Shared Focus Fuels Product Success

Speaker: Anne Steiner and David Laribee

As a concept, Developer Experience (DX) has gained significant attention in the tech industry. It emphasizes engineers’ efficiency and satisfaction during the product development process. As product managers, we need to understand how a good DX can contribute not only to the well-being of our development teams but also to the broader objectives of product success and customer satisfaction.

article thumbnail

Our Top 5 Data Mesh Articles In 2022

Monte Carlo

Data mesh is a complex socio-technological data engineering concept, but it doesn’t change too much. The four principles are still the four principles, there are still three experience planes, and automation is still as vital as ever. This is a good thing! Data mesh is one of those rare transformative concepts that emerged relatively fully formed as a result of creator Zhamak Dhegani’s years of consulting experience captured in a comprehensive 384 page book.

Retail 52
article thumbnail

Increase Your Odds Of Success For Analytics And AI Through More Effective Knowledge Management With AlignAI

Data Engineering Podcast

Summary Making effective use of data requires proper context around the information that is being used. As the size and complexity of your organization increases the difficulty of ensuring that everyone has the necessary knowledge about how to get their work done scales exponentially. Wikis and intranets are a common way to attempt to solve this problem, but they are frequently ineffective.

article thumbnail

Everything Best Of Analytics for 2023: 7 Must Read Articles!

U-Next

Introduction . If you have access to data regarding every aspect of the business you work for, then you are sitting on a goldmine. Data is the most crucial, important, and valued asset of today’s technology. Every emerging technology – Artificial Intelligence, Cloud Computing, Cybersecurity, Machine Learning, etc., are all dependent on data. They either work towards extracting, storing, or protecting data, making it one of the most priceless assets an organization could own.

Food 40
article thumbnail

A Guide to Train an Image Classification Model Using Tensorflow

KDnuggets

Classify images at scale and with very high accuracy with the advent of machine learning and deep learning algorithms.

article thumbnail

How to Build an Experimentation Culture for Data-Driven Product Development

Speaker: Margaret-Ann Seger, Head of Product, Statsig

Experimentation is often seen as an aspirational practice, especially at smaller, fast-moving companies who are strapped for time and resources. So, how can you get your team making decisions in a more data-driven way while continuing to remain lean and maintaining ship velocity? In this webinar, Margaret-Ann Seger, Head of Product at Statsig, will teach you how to build an experimentation culture from the ground-up, graduating from just getting started with data-driven development to operating