#1 Tech Interview Platform

Data Mesh is a decentralized approach to data architecture where, instead of one central team owning a monolithic lake/warehouse, individual domain teams own and serve their own data as a product. Parquet, ORC, Avro, CSV, and JSON are commonly used file formats in big data and data engineering. An index is a data structure typically a B-tree, or a hash index for exact-match lookups that allows a database to find rows without scanning the entire table. Slowly Changing Dimensions define how a data warehouse handles changes to dimension attributes over time for example, when a customer’s address or a product’s category changes. Repartition() and coalesce() are Spark operations used to change the number of partitions in an RDD, DataFrame, or Dataset.
Modern samplers like py-spy collect stack snapshots via ptrace without GIL interference, suiting production profiling. For blocking work shunted to threads (run_in_executor), contextvars propagate automatically from Python 3.11 onward; earlier versions require contextvars.copy_context().run(fn). Each ContextVar has a default; Context.run() or var.set() returns a token for restoration, forming an immutable snapshot.

  • Views in SQL are a kind of virtual table.
  • In summary, ETL processes extract data from multiple sources, transform it into a suitable format, and load it into a data warehouse for combined historical and current data analysis.
  • Most strong frameworks combine getByRole first, with getByTestId as a fallback for elements that have no clear semantic role.
  • It has an interactive PySpark shell for analyzing structured and semi-structured data in a distributed setting.

Always check the status_code member variable on the response object and handle things according to the code. NumPy is a library for numerical computations, that can contains large, multi dimensional arrays and matrices with mathematical functions. The __slots__ feature limits which attributes can be added to instances of a class to save memory, since a dict is not created for each instance The metaclass overrides the __call__ method to control instance creation, storing a single instance in a class-level dictionary. Magic methods (dunder methods) are special methods with double underscores, allowing customization of object behaviour for built-in operations. Classes are instances of a metaclass, and metaclasses can define a method of creating classes that can, for instance, alter class attributes or constrain the class to a particular design pattern.

Q2. What is the difference between groupBy and reduceByKey in PySpark?

The NameNode uses these reports to maintain an accurate mapping of files to blocks and their replicas. Apache Hadoop is widely used for processing and storing massive datasets because of its distributed architecture. DataNodes handle read and write requests from clients and report their status to the NameNode.

What you will need to get started

AQE detects skewed partitions and splits them into smaller tasks automatically. The uneven distribution of product categories causes some tasks to run much longer than others, creating a bottleneck. Without all three, you risk “at-least-once” delivery (possible duplicates on failure recovery). Structured Streaming achieves this through a combination of checkpointing, idempotent sinks, and offset tracking.

Additionally, FastAPI is gaining traction for its high-performance capabilities, utilizing Python’s asynchronous features for speed. Leverage stored procedures in SQL for complex operations that can enhance performance. https://uvik.io/ When integrating Python with SQL databases, it’s crucial to follow best practices to ensure seamless functionality.