CertsGate
See all results for ""
Home Exams
CRISC ISACA CISSP ISC2 200-301 Cisco SY0-701 CompTIA AZ-104 Microsoft AI-900 Microsoft AIGP IAPP 1Z0-1067-26 Oracle View All Exams →
Sign in Create account

Databricks Certified Data Engineer Associate Databricks-Certified-Data-Engineer-Associate Exam Questions

Preparing for the Databricks-Certified-Data-Engineer-Associate exam is simple with CertsGate. We offer easy-to-understand study materials that help you learn the most important exam topics. You can study using our PDF questions, practice online with a real exam-style test, or use the desktop practice software. Choose the study method that works best for you and prepare at your own pace.

At CertsGate, we keep our Databricks-Certified-Data-Engineer-Associate practice questions up to date. Whenever the exam syllabus or objectives change, we update our study materials so you always learn the latest topics. This helps you save time, avoid outdated content, and feel more confident when you take your exam.

Download Exam View Entire Exam
Page: 1 / 1
Question #1 (Topic: Demo Questions)

A new data engineering team team has been assigned to an ELT project. The new data engineering team will need full privileges on the table sales to fully manage the project.

Which command can be used to grant full permissions on the database to the new data engineering team?

A.

grant all privileges on table sales TO team;

B.

GRANT SELECT ON TABLE sales TO team;

C.

GRANT SELECT CREATE MODIFY ON TABLE sales TO team;

D.

GRANT ALL PRIVILEGES ON TABLE team TO sales;

Correct Answer: A
Explanation:

To grant full privileges on a table such as ' sales ' to a group like ' team ' , the correct SQL command in Databricks is:

GRANT ALL PRIVILEGES ON TABLE sales TO team;

This command assigns all available privileges, including SELECT, INSERT, UPDATE, DELETE, and any other data manipulation or definition actions, to the specified team. This is typically necessary when a team needs full control over a table to manage and manipulate it as part of a project or ongoing maintenance.

[References:Databricks documentation on SQL permissions: SQL Permissions in Databricks, , , , , ]
Question #2 (Topic: Demo Questions)

Which TWO items are characteristics of the Gold Layer?

Choose 2 answers

A.

Read-optimized

B.

Normalised

C.

Raw Data

D.

Historical lineage

E.

De-normalised

Correct Answer: A, E
Question #3 (Topic: Demo Questions)

A data engineer converts an external Delta table to a Unity Catalog managed table. A Structured Streaming job that reads from the table continues running during the conversion. After the conversion completes, the streaming job stops processing new records.

How should the data engineer resolve the issue?

A.

Restart the streaming job so that it uses the new managed-table location.

B.

Run REFRESH TABLE on the converted table to update the streaming checkpoint.

C.

Grant the streaming job additional permissions on the new managed-storage location.

D.

Delete the streaming checkpoint directory and reprocess the complete source from the beginning.

Correct Answer: A
Explanation:

Databricks intentionally stops existing streaming reads and writes after an external Delta table is converted to a Unity Catalog managed table. This protects data consistency because the table’s managed storage location and access path have changed. The documented recovery action is to restart the stream with the same configuration. The existing checkpoint can then resume processing from the last committed offset, and supported path-based access is redirected to the converted managed table. REFRESH TABLE invalidates cached metadata but does not repair a stopped Structured Streaming query or modify its checkpoint. Deleting the checkpoint would discard progress information and could cause unnecessary reprocessing or duplicate output. Additional permissions are not the indicated solution when the failure is specifically caused by managed-table conversion. Therefore, option A is correct.

Question #4 (Topic: Demo Questions)

A data engineer has developed a data pipeline to ingest data from a JSON source using Auto Loader, but the engineer has not provided any type inference or schema hints in their pipeline. Upon reviewing the data, the data engineer has noticed that all of the columns in the target table are of the string type despite some of the fields only including float or boolean values.

Which of the following describes why Auto Loader inferred all of the columns to be of the string type?

A.

There was a type mismatch between the specific schema and the inferred schema

B.

JSON data is a text-based format

C.

Auto Loader only works with string data

D.

All of the fields had at least one null value

E.

Auto Loader cannot infer the schema of ingested data

Correct Answer: B
Explanation:

JSON data is a text-based format that represents data as a collection of name-value pairs. By default, when Auto Loader infers the schema of JSON data, it treats all columns as strings. This is because JSON data can have varying data types for the same column across different files or records, and Auto Loader does not attempt to reconcile these differences. For example, a column named “age” may have integer values in some files, but string values in others. To avoid data loss or errors, Auto Loader infers the column as a string type. However, Auto Loader also provides an option to infer more precise column types based on the sample data. This option is called cloudFiles.inferColumnTypes and it can be set to true or false. When set to true, Auto Loader tries to infer the exact data types of the columns, such as integers, floats, booleans, or nested structures. When set to false, Auto Loader infers all columns as strings. The default value of this option is false. References: Configure schema inference and evolution in Auto Loader, Schema inference with auto loader (non-DLT and DLT), Using and Abusing Auto Loader’s Inferred Schema, Explicit path to data or a defined schema required for Auto loader.

Question #5 (Topic: Demo Questions)

A data engineer needs to optimize the data layout and query performance for an e-commerce transactions Delta table. The table is partitioned by " purchase_date " a date column which helps with time-based queries but does not optimize searches on user statistics " customer_id " , a high-cardinality column.

The table is usually queried with filters on " customer_i

d " within specific date ranges, but since this data is spread across multiple files in each partition, it results in full partition scans and increased runtime and costs.

How should the data engineer optimize the Data Layout for efficient reads?

A.

Alter table implementing liquid clustering on " customerid " while keeping the existing partitioning.

B.

Alter the table to partition by " customer_id " .

C.

Enable delta caching on the cluster so that frequent reads are cached for performance.

D.

Alter the table implementing liquid clustering by " customer_id " and " purchase_date " .

Correct Answer: D
Download Exam
Page: 1 / 1
Next Page