← Back to data subTLDR

data subTLDR week 36 year 2026

r/MachineLearningr/dataengineeringr/SQL

Unveiling the Humor in SQL Naming, Tackling Tough SQL Interview Questions, Striking Balance in 'Clever' SQL Queries, Embracing AI with Cautious Optimism, Exploring Apache Airflow's Potential in 'Beyond the Dag' Hackathon

Week 36, 2026
Posted in r/MachineLearningbyu/DataShack9/2/2026
880

I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]

Project
A developer has collected and uploaded a comprehensive dataset of 5.94 billion TikTok videos and 3.23 billion profiles to Hugging Face, using a reverse-engineering method on the TikTok mobile app. This public data, which includes profiles, comments, hashtags, and more, was accessed without a TikTok account, though some speculate it might be against the platform's ToS. The developer also offers full documentation free of charge, but the complete codebase comes at a price. The news is generally met with enthusiasm, with some users expressing interest in downloading the dataset immediately, though a few voiced concerns about the legality of charging for this data.
201 comments
Share
Save
View on Reddit →
Posted in r/SQLbyu/eww19919/2/2026
514

Ah yes, the table lake

Discussion
The discussion centers around the humorous and risky aspects of naming variables in programming, particularly in SQL. Participants stress the importance of properly parameterizing queries to avoid potential issues, illustrated through a joke about inadvertently dropping a data lake. There's also a reference to the infamous Bobby Tables XKCD comic, which highlights the dangers of SQL injection. The overall sentiment is positive, with participants exchanging jokes and advice about safe coding practices.
19 comments
Share
Save
View on Reddit →
Posted in r/dataengineeringbyu/informatica69/4/2026
325

AI is really freaking me out

Career
While AI's rapid advancement is causing concern among some data engineers, others argue the field is not disappearing but merely evolving. They note that while AI may automate coding tasks, it can't replace the problem-solving, decision-making, and monitoring responsibilities of engineers. AI's efficiency is acknowledged, but its limitations are also highlighted, particularly its tendency to make mistakes if not constantly supervised. Some see opportunities in fixing AI-generated errors and developing interfaces for AI agents. The consensus seems to be that while job roles might change due to AI, the need for human expertise remains essential. Sentiments are mixed but lean towards cautious optimism.
195 comments
Share
Save
View on Reddit →
Posted in r/MachineLearningbyu/tariban8/31/2026
273

Cold emailing profs about PhD positions? Read this [D]

Discussion
When cold emailing professors about PhD positions, it's crucial to be concise, specific about your interests, and respectful of instructions on their website. Many professors appreciate meeting prospective students at conferences and value authenticity over attempts to mirror their interests. Checking current grad students' work can provide insight into the lab's focus. It's effective to convey how you could build upon the professor's work rather than summarizing it. Avoid generic interests and excessive AI use. Dishonesty, like passing off workshop papers as conference papers, is a red flag. Overall, the sentiment is positive toward thoughtfully approached cold emailing.
31 comments
Share
Save
View on Reddit →
Posted in r/SQLbyu/Temporary-Cup-21409/4/2026
71

What was the toughest SQL interview question you have faced so far?

MySQL
The toughest SQL interview questions often test not just technical knowledge, but also how candidates handle ambiguity and unexpected challenges. Some interviewees found open-ended questions about optimizing slow stored procedures challenging but insightful. Others were stumped by questions with multiple correct answers or ones requiring written or memorized responses. A recurring theme was the difficulty of questions that didn't match the role's advertised requirements. Toxicity, such as an interviewer reacting angrily to a valid solution, was a red flag for candidates. Overall, the thread sentiment was mixed, with some seeing the value in tough questions, while others found them frustrating or misleading.
62 comments
Share
Save
View on Reddit →
Posted in r/SQLbyu/No_Ambition83239/3/2026
49

When does a SQL query become “too clever”?

Discussion
SQL queries become too clever when they are difficult to fix, take longer to read, or obfuscate the purpose. While some developers prioritize technical efficiency, many argue that maintainability is key. This includes breaking complex queries into simpler, manageable parts and providing useful comments. The number of lines is considered less important than maintainability and query performance. However, all agree that a well-structured, easily understandable query is preferable over a compact, complex one that may cause difficulties in troubleshooting and optimization later. The consensus leans towards ensuring code is not over-engineered at the expense of speed and maintainability. The overall sentiment is mixed, leaning towards caution against over-complication.
64 comments
Share
Save
View on Reddit →
Posted in r/dataengineeringbyu/prenomenon8/31/2026
45

Beyond the Dag - a data engineering hackathon

Open Source
The 'Beyond the Dag' data engineering hackathon is focused on Apache Airflow, emphasizing its evolution beyond data movement and encouraging participants to explore its potential. The competition is seen as a unique opportunity to enhance data engineering profiles due to the scarcity of similar hackathons. The discussion highlighted the potential of AI agents in Airflow, leading to a detailed explanation of how to use them and the advantages they offer. However, there was also concern about the rapid pace of technology change and the challenges of migrating from one version of Airflow to another. The overall sentiment was mixed but leaned towards curiosity and interest.
7 comments
Share
Save
View on Reddit →
Posted in r/dataengineeringbyu/Comprehensive_Level79/1/2026
42

extracting data from on-prem databases (sql server, oracle, postgres) to cloud, which options you guys recommend?

Discussion
The thread discusses the challenges of data migration from on-premises databases to the cloud, with a focus on Azure Data Factory (ADF) and alternatives. The sentiment is mixed, with some users suggesting the need for a data architect and project manager to ensure secure and efficient migration. Others argue that if ADF is working reliably, it may not be worth changing. Alternatives such as AWS DMS, Airbyte, Debezium, and local Airflow are mentioned, but there are concerns about the effort and complexity of managing these solutions. Some users recommend keeping ADF for extraction and using other tools for orchestration and transformation. They also emphasize the importance of considering the client's comfort level and technical capabilities.
46 comments
Share
Save
View on Reddit →

Subscribe to data-subtldr

Get weekly summaries of top content from r/dataengineering, r/MachineLearning and more directly in your inbox.

Get the weekly data subTLDR in your inbox!

We respect your privacy. No spam, ever.