chapter
    Processor Performance, Pipelining and Hazards PYQs for GATE CS

    Solve 13+ Processor Performance, Pipelining and Hazards previous year questions for GATE CS with answers and detailed solutions. Free sample questions below.

    Try a question

    Answer it here to see how it works. Nothing is recorded until you sign in.

    Question 1
    2026 Slot Set2 PYQ
    Level 3: Exam Standard
    A non-pipelined instruction execution unit that operates at 1.6 GHz clock takes an average of 5 clock cycles to complete the execution of an instruction. To improve the performance, the system was pipelined with a goal of achieving an average throughput of one instruction per clock cycle. However, it could operate only at 1.2 GHz due to pipeline overheads. While executing a program in the pipelined design, of instructions encountered a stall of 2 cycles due to pipeline hazards. The speed-up obtained by the pipelined design over the non-pipelined one for this program is ___________. (rounded off to two decimal places)

    Note:
    Question 2
    2026 Slot Set1 PYQ
    Level 3: Exam Standard
    The EX stage of a pipelined processor performs the memory read operations for LOAD instructions, and the operations for the arithmetic and logic instructions. Let denote the time taken by the EX stage to perform the operation for an instruction. For each instruction type, the values of and M (the number of instructions of that type in a sequence of 100 instructions for a program P), are given in the table below.

    The duration of the pipeline clock cycle is 1 nanosecond. Assume that the latch time for the interstage buffers in the pipeline is negligible.

    Instruction in
    nanoseconds
    LOAD1.815
    IMUL1.510
    IDIV2.55
    FADD1.710
    FSUB1.75
    FMUL2.815
    FDIV3.25
    All other
    instructions
    Less than
    1.0
    35


    When program P is executed, the number of clock cycles for which the pipeline is stalled due to structural hazards in the EX stage is ______. (answer in integer)
    Question 3
    2026 Slot Set1 PYQ
    Level 3: Exam Standard

    Which one of the following dependencies among the register operands of different instructions can cause a data hazard in a pipelined processor?

    Question 4
    2025 Slot Set2 PYQ
    Level 3: Exam Standard
    An application executes number of instructions in 6.3 seconds. There are four types of instructions, the details of which are given in the table. The duration of a clock cycle in nanoseconds is _________. (rounded off to one decimal place)

    Instruction typeClock cycles required perinstruction (CPI)Number of instructionsexecutedBranch22.25 × 10⁸Load51.20 × 10⁸Store41.65 × 10⁸Arithmetic31.30 × 10⁸
    Question 5
    2025 Slot Set2 PYQ
    Level 3: Exam Standard

    A 5-stage instruction pipeline has stage delays of 180, 250, 150, 170, and 250, respectively, in nanoseconds. The delay of an inter-stage latch is 10 nanoseconds. Assume that there are no pipeline stalls due to branches and other hazards. The time taken to process 1000 instructions in microseconds is __________ . (rounded off to two decimal places)

    Question 6
    2024 Slot Set2 PYQ
    Level 3: Exam Standard

    A non-pipelined instruction execution unit operating at 2 GHz takes an average of 6 cycles to execute an instruction of a program P. The unit is then redesigned to operate on a 5-stage pipeline at 2 GHz. Assume that the ideal throughput of the pipelined unit is 1 instruction per cycle. In the execution of program P, 20% instructions incur an average of 2 cycles stall due to data hazards and 20% instructions incur an average of 3 cycles stall due to control hazards. The speedup <i>(rounded off to one decimal place)</i> obtained by the pipelined design over the non-pipelined design is __________

    Question 7
    2024 Slot Set2 PYQ
    Level 3: Exam Standard
    An instruction format has the following structure:

    Instruction Number: Opcode destination reg, source reg-1, source reg-2

    Consider the following sequence of instructions to be executed in a pipelined processor:

    I1: DIV R3, R1, R2
    I2: SUB R5, R3, R4
    I3: ADD R3, R5, R6
    I4: MUL R7, R3, R8

    Which of the following statements is/are TRUE?
    Question 8
    2024 Slot Set1 PYQ
    Level 3: Exam Standard
    The baseline execution time of a program on a 2 GHz single core machine is 100 nanoseconds (). The code corresponding to 90% of the execution time can be fully parallelized. The overhead for using an additional core is 10 when running on a multicore system. Assume that all cores in the multicore system run their share of the parallelized code for an equal amount of time.

    The number of cores that minimize the execution time of the program is _________
    Question 9
    2024 Slot Set1 PYQ
    Level 3: Exam Standard

    Consider a 5-stage pipelined processor with Instruction Fetch (IF), Instruction Decode (ID), Execute (EX), Memory Access (MEM), and Register Writeback (WB) stages. Which of the following statements about forwarding is/are CORRECT?

    Question 10
    2023 PYQ
    Level 3: Exam Standard
    Consider a 3-stage pipelined processor having a delay of 10 ns (nanoseconds), 20 ns, and 14 ns, for the first, second, and the third stages, respectively. Assume that there is no other delay and the processor does not suffer from any pipeline hazards. Also assume that one instruction is fetched every cycle.

    The total execution time for executing 100 instructions on this processor is __________ ns.
    Free preview ends here

    Login to view the complete previous-year questions and solutions

    Creating an account is free. You get the rest of this chapter, step-by-step solutions, and a study plan built around the topics you are actually weak at.

    Why MastersUp

    Personalised first. High quality throughout.

    Most platforms hand everyone the same content. Here the content moves with your performance, topic by topic.

    Built around you, not around a syllabus PDF

    Every answer you give moves your topic-level intelligence rate. The next question, the next revision card and tomorrow's plan all change with it.

    Revision that hits your weak spots

    We only revise topics you have actually attempted and are still below the safe bar on — never the same chapter on repeat.

    Questions calibrated to the real exam

    Each question carries a measured toughness. You are served a rung above your current level, so practice keeps stretching you.

    Notes written for recall, not for volume

    Full lesson cards for first study, curated short-note cards for the last mile — with derivations, traps and exam patterns marked.

    One place for everything

    Notes, chapter practice, previous-year questions, test series and full-length papers — all feeding one picture of your preparation.

    Honest progress

    No vanity streaks. Progress here means chapters mastered and accuracy that held up on harder questions.

    Unlock the whole course

    Full notes and short notes, the complete question bank with worked solutions, mock tests, full-length papers, and an adaptive plan that rebuilds itself as you improve.

    Processor Performance, Pipelining and Hazards PYQs for GATE CS

    Solve 13+ Processor Performance, Pipelining and Hazards previous year questions for GATE CS with answers and detailed solutions. Free sample questions below.

    Chapter Roadmap: Processor Performance, Pipelining and Hazards

    Chapter Roadmap

    Your journey through Processor Performance and Pipelining. Master these five stages to decode any pipeline question.

    1. Pipeline Timing, Latency and Throughput

    t1

    The foundation. How overlapping execution changes cycle time, total latency, and overall throughput.

    Weightage: Foundational / Moderate

    2. Pipeline CPI, Clock Rate and Stall Performance

    t2

    Moving beyond ideal scenarios. Calculating the real Cycles Per Instruction when stalls occur.

    Weightage: High

    3. Data Dependencies, Hazards and Forwarding

    t3

    Identifying Read-After-Write, Write-After-Read, and Write-After-Write hazards, and resolving them with data forwarding.

    Weightage: High

    4. Control Hazards and Branch Prediction

    t4

    Handling jumps and branches. Static prediction, dynamic prediction, and branch target buffers.

    Weightage: Moderate

    5. Parallel Speedup and Multicore Performance

    t5

    Amdahl's Law, Gustafson's Law, and the limits of scaling across multiple cores.

    Weightage: Low to Moderate

    Topic Hero: Pipeline Timing, Latency and Throughput

    Pipeline Timing, Latency and Throughput

    Overlapping execution to maximize instruction completion rate.

    The Core Insight: Pipelining does not reduce the time it takes to execute a single instruction (latency). Instead, it increases the number of instructions completed per second (throughput) by overlapping their execution stages, much like an assembly line.

    Latency Time for 1 instruction
    Throughput Instructions per unit time

    Processor Performance, Pipelining and Hazards: Solved Questions with Step-by-Step Explanations (10 Problems)

    Question 1 · Computer Organization and Architecture · 2026_Set2 NAT
    A non-pipelined instruction execution unit that operates at 1.6 GHz clock takes an average of 5 clock cycles to complete the execution of an instruction. To improve the performance, the system was pipelined with a goal of achieving an average throughput of one instruction per clock cycle. However, it could operate only at 1.2 GHz due to pipeline overheads. While executing a program in the pipelined design, of instructions encountered a stall of 2 cycles due to pipeline hazards. The speed-up obtained by the pipelined design over the non-pipelined one for this program is ___________. (rounded off to two decimal places)

    Note:
    Correct Answer:

    2.34

    Step-by-Step Solution

    Key idea: This is a pipelining speedup NAT with clock rate change, recognisable because it gives the clock frequencies and CPI/stall information for both non-pipelined and pipelined designs.

    Step 1: Calculate the execution time per instruction for the non-pipelined design.

    Clock frequency GHz.

    Cycles per instruction .

    Time per instruction time units.

    Step 2: Calculate the execution time per instruction for the pipelined design.

    Clock frequency GHz.

    Base CPI = 1.

    Stall penalty: 30% of instructions encounter a 2-cycle stall.

    Average stall cycles per instruction = .

    Actual CPI .

    Time per instruction time units.

    Step 3: Calculate the speedup.

    Speedup = .

    Rounded to two decimal places, this is 2.34.

    Answer: 2.34

    Question 2 · Computer Organization and Architecture · 2026_Set1 NAT
    The EX stage of a pipelined processor performs the memory read operations for LOAD instructions, and the operations for the arithmetic and logic instructions. Let denote the time taken by the EX stage to perform the operation for an instruction. For each instruction type, the values of and M (the number of instructions of that type in a sequence of 100 instructions for a program P), are given in the table below.

    The duration of the pipeline clock cycle is 1 nanosecond. Assume that the latch time for the interstage buffers in the pipeline is negligible.

    Instruction in
    nanoseconds
    LOAD1.815
    IMUL1.510
    IDIV2.55
    FADD1.710
    FSUB1.75
    FMUL2.815
    FDIV3.25
    All other
    instructions
    Less than
    1.0
    35


    When program P is executed, the number of clock cycles for which the pipeline is stalled due to structural hazards in the EX stage is ______. (answer in integer)
    Correct Answer:

    95

    Step-by-Step Solution

    Key idea: This is a structural hazard stall calculation NAT, recognisable because it gives varying execution times for the EX stage and asks for the total number of stall cycles.

    Step 1: Understand the clock cycle constraint. The pipeline clock cycle is 1 ns. Any instruction whose EX stage takes longer than 1 ns must occupy the EX stage for multiple clock cycles.

    Step 2: Calculate the number of clock cycles each instruction type spends in the EX stage. This is given by .

    LOAD: cycles.

    IMUL: cycles.

    IDIV: cycles.

    FADD: cycles.

    FSUB: cycles.

    FMUL: cycles.

    FDIV: cycles.

    Other: cycle.

    Step 3: Calculate the number of stall cycles caused by each instruction. An instruction occupying the EX stage for cycles causes stall cycles for subsequent instructions.

    LOAD: stall cycle. Total for 15 instructions = .

    IMUL: . Total = .

    IDIV: . Total = .

    FADD: . Total = .

    FSUB: . Total = .

    FMUL: . Total = .

    FDIV: . Total = .

    Other: . Total = 0.

    Step 4: Sum the total stall cycles.

    Total stalls = .

    Answer: 95

    Question 3 · Computer Organization and Architecture · 2026_Set1 MCQ

    Which one of the following dependencies among the register operands of different instructions can cause a data hazard in a pipelined processor?

    1. A.

      Read-after-read

    2. B.

      Read-after-write

    3. C.

      Write-after-read

    4. D.

      Write-after-write

    Correct Answer:

    B

    Step-by-Step Solution

    Key idea: This is a conceptual MCQ on data dependency types, recognisable because it asks which dependency among register operands causes a data hazard in a pipelined processor.

    Step 1: Recall the three types of data dependencies between instructions I (earlier) and J (later):

    • RAW (Read After Write): J reads a register that I writes. This is a true data dependency because J genuinely needs the value I produces.
    • WAR (Write After Read): J writes a register that I reads. This is a name dependency (anti-dependency).
    • WAW (Write After Write): J writes a register that I also writes. This is a name dependency (output dependency).

    Step 2: In a standard in-order 5-stage pipeline, instructions execute in program order. WAR and WAW hazards cannot occur because reads always happen before later writes in program order, and writes complete in order. Only RAW creates a genuine hazard where J might read a stale value before I has written the new one.

    Step 3: Read-after-read (option A) is not even a recognized dependency type that causes any hazard, since both instructions only read and neither modifies the value.

    Step 4: Therefore, only Read-after-write (RAW) causes a data hazard in a pipelined processor.

    Answer: B

    Question 4 · Computer Organization and Architecture · 2025_Set2 NAT
    An application executes number of instructions in 6.3 seconds. There are four types of instructions, the details of which are given in the table. The duration of a clock cycle in nanoseconds is _________. (rounded off to one decimal place)

    Instruction typeClock cycles required perinstruction (CPI)Number of instructionsexecutedBranch22.25 × 10⁸Load51.20 × 10⁸Store41.65 × 10⁸Arithmetic31.30 × 10⁸
    Correct Answer:

    3.0

    Step-by-Step Solution

    Insight: The total execution time is the product of total cycles and the clock cycle duration. We can find total cycles by summing the cycles for each instruction type.

    Exam route: Total cycles = (2.25×2 + 1.20×5 + 1.65×4 + 1.30×3) × 10⁸ = (4.5 + 6.0 + 6.6 + 3.9) × 10⁸ = 21.0 × 10⁸ cycles. Total time = 6.3 s = 63 × 10⁸ ns. Cycle duration = 63 / 21 = 3.0 ns.

    Learning route:

    1. Identify the number of cycles for each instruction type by multiplying its count by its CPI.
    2. Sum these to get the total cycles executed: 4.5e8 + 6.0e8 + 6.6e8 + 3.9e8 = 21.0e8 cycles.
    3. Convert the total execution time to nanoseconds: 6.3 seconds = 6.3 × 10⁹ ns = 63 × 10⁸ ns.
    4. Divide the total time in nanoseconds by the total number of cycles to find the duration of a single clock cycle: (63 × 10⁸) / (21 × 10⁸) = 3.0 ns.
    Question 5 · Computer Organization and Architecture · 2025_Set2 NAT

    A 5-stage instruction pipeline has stage delays of 180, 250, 150, 170, and 250, respectively, in nanoseconds. The delay of an inter-stage latch is 10 nanoseconds. Assume that there are no pipeline stalls due to branches and other hazards. The time taken to process 1000 instructions in microseconds is __________ . (rounded off to two decimal places)

    Correct Answer:

    261.04

    Step-by-Step Solution

    Key idea: This is a pipeline timing NAT, recognisable because it provides stage delays, latch delays, and an instruction count, asking for the total execution time.

    Step 1: Determine the pipeline clock cycle time. The cycle time is governed by the slowest stage plus the inter-stage latch delay.

    Max stage delay = ns.

    Latch delay = 10 ns.

    Cycle time ns.

    Step 2: Calculate the total number of clock cycles for 1000 instructions.

    For a -stage pipeline processing instructions with no stalls, the total number of cycles is .

    Here, and .

    Total cycles = cycles.

    Step 3: Calculate the total execution time.

    Total time = ns = ns.

    Step 4: Convert to microseconds.

    .

    Answer: 261.04

    Question 6 · Computer Organization and Architecture · 2024_Set2 NAT

    A non-pipelined instruction execution unit operating at 2 GHz takes an average of 6 cycles to execute an instruction of a program P. The unit is then redesigned to operate on a 5-stage pipeline at 2 GHz. Assume that the ideal throughput of the pipelined unit is 1 instruction per cycle. In the execution of program P, 20% instructions incur an average of 2 cycles stall due to data hazards and 20% instructions incur an average of 3 cycles stall due to control hazards. The speedup <i>(rounded off to one decimal place)</i> obtained by the pipelined design over the non-pipelined design is __________

    Correct Answer:

    3.0

    Step-by-Step Solution

    Insight: Speedup is the ratio of non-pipelined execution time to pipelined execution time. Both depend on their respective CPIs and the clock cycle time.

    Exam route: Non-pipe time/inst = 6 × (1/2 GHz) = 3 ns. Pipe avg stall = 0.2×2 + 0.2×3 = 1.0. Pipe CPI = 1 + 1.0 = 2.0. Pipe time/inst = 2.0 × (1/2 GHz) = 1 ns. Speedup = 3 / 1 = 3.0.

    Learning route:

    1. Calculate the non-pipelined time per instruction: CPI_non_pipe × Cycle_time = 6 × (1 / 2×10⁹) s = 3 ns.
    2. Calculate the average stall cycles per instruction for the pipelined design: (0.20 × 2) + (0.20 × 3) = 0.4 + 0.6 = 1.0 cycle.
    3. Calculate the actual pipelined CPI: Ideal CPI + Average stalls = 1 + 1.0 = 2.0.
    4. Calculate the pipelined time per instruction: CPI_pipe × Cycle_time = 2.0 × (1 / 2×10⁹) s = 1 ns.
    5. Compute the speedup: Time_non_pipe / Time_pipe = 3 ns / 1 ns = 3.0.
    Question 7 · Computer Organization and Architecture · 2024_Set2 MSQ
    An instruction format has the following structure:

    Instruction Number: Opcode destination reg, source reg-1, source reg-2

    Consider the following sequence of instructions to be executed in a pipelined processor:

    I1: DIV R3, R1, R2
    I2: SUB R5, R3, R4
    I3: ADD R3, R5, R6
    I4: MUL R7, R3, R8

    Which of the following statements is/are TRUE?
    1. A.

      There is a RAW dependency on R3 between I1 and I2

    2. B.

      There is a WAR dependency on R3 between I1 and I3

    3. C.

      There is a RAW dependency on R3 between I2 and I3

    4. D.

      There is a WAW dependency on R3 between I3 and I4

    Correct Answer:

    ["A"]

    Step-by-Step Solution

    Key idea: This is an instruction-level dependency analysis MSQ, recognisable because it gives a sequence of instructions with a specific format and asks which dependency statements are true.

    Step 1: Parse the instruction format carefully. The format is: Opcode destination, source-1, source-2. So the first register after the opcode is the destination (written), and the next two are sources (read).

    Step 2: List reads and writes for each instruction:

    • I1: DIV R3, R1, R2 W: R3; R: R1, R2
    • I2: SUB R5, R3, R4 W: R5; R: R3, R4
    • I3: ADD R3, R5, R6 W: R3; R: R5, R6
    • I4: MUL R7, R3, R8 W: R7; R: R3, R8

    Step 3: Evaluate option A: RAW on R3 between I1 and I2. I1 writes R3, I2 reads R3. This is Read After Write on R3. TRUE.

    Step 4: Evaluate option B: WAR on R3 between I1 and I3. For WAR, the earlier instruction must read R3 and the later one must write R3. I1 writes R3 (does not read it), I3 writes R3. Both write R3, so this is WAW, not WAR. FALSE.

    Step 5: Evaluate option C: RAW on R3 between I2 and I3. For RAW, I2 must write R3 and I3 must read R3. I2 writes R5 (not R3), and I3 writes R3 (does not read it). I2 reads R3 and I3 writes R3, which is WAR, not RAW. FALSE.

    Step 6: Evaluate option D: WAW on R3 between I3 and I4. For WAW, both must write R3. I3 writes R3, but I4 writes R7 (and reads R3). I3 writes R3 and I4 reads R3, which is RAW, not WAW. FALSE.

    Answer: A

    Question 8 · Computer Organization and Architecture · 2024_Set1 NAT
    The baseline execution time of a program on a 2 GHz single core machine is 100 nanoseconds (). The code corresponding to 90% of the execution time can be fully parallelized. The overhead for using an additional core is 10 when running on a multicore system. Assume that all cores in the multicore system run their share of the parallelized code for an equal amount of time.

    The number of cores that minimize the execution time of the program is _________
    Correct Answer:

    3

    Step-by-Step Solution

    Key idea: This is a multicore optimization NAT, recognisable because it gives baseline execution time, a parallelizable fraction, and per-core overhead, asking for the number of cores that minimizes total execution time.

    Step 1: Decompose the baseline execution time into sequential and parallel parts. Baseline ns. The parallelizable fraction is , so parallel time ns. Sequential time ns.

    Step 2: Model the total execution time with cores. The sequential part runs on one core and takes 10 ns regardless. The parallel part is divided equally among cores, taking ns. The overhead for using additional cores is ns per additional core, i.e., ns.

    Step 3: Write the total time function: .

    Step 4: Minimize by taking the derivative with respect to and setting it to zero: .

    Step 5: Solve: (taking the positive root since must be positive).

    Step 6: Verify by checking nearby integer values: ns, ns, ns. The minimum is at .

    Answer: 3

    Question 9 · Computer Organization and Architecture · 2024_Set1 MSQ

    Consider a 5-stage pipelined processor with Instruction Fetch (IF), Instruction Decode (ID), Execute (EX), Memory Access (MEM), and Register Writeback (WB) stages. Which of the following statements about forwarding is/are CORRECT?

    1. A.

      In a pipelined execution, forwarding means the result from a source stage of an earlier instruction is passed on to the destination stage of a later instruction

    2. B.

      In forwarding, data from the output of the MEM stage can be passed on to the input of the EX stage of the next instruction

    3. C.

      Forwarding cannot prevent all pipeline stalls

    4. D.

      Forwarding does not require any extra hardware to retrieve the data from the pipeline stages

    Correct Answer:

    ["A","B","C"]

    Step-by-Step Solution

    Key idea: This is a conceptual MSQ on data forwarding in pipelines, recognisable because it asks which statements about forwarding are correct in a standard 5-stage pipeline.

    Step 1: Evaluate option A. Forwarding (bypassing) means taking the result produced at a later pipeline stage (e.g., EX or MEM output) of an earlier instruction and routing it directly to an earlier stage (e.g., EX input) of a later instruction, avoiding the wait for register writeback. This is the standard definition. Option A is CORRECT.

    Step 2: Evaluate option B. In a 5-stage pipeline, the MEM stage output holds the result of the previous instruction. This data can be forwarded to the EX stage input of the next instruction via a MEM-to-EX forwarding path. This is a standard forwarding path. Option B is CORRECT.

    Step 3: Evaluate option C. Forwarding resolves most RAW hazards, but it cannot resolve the load-use hazard. When a load instruction is immediately followed by an instruction that uses the loaded value, the data is only available at the end of the MEM stage, but the next instruction needs it at the beginning of EX. A 1-cycle stall is unavoidable even with forwarding. Option C is CORRECT.

    Step 4: Evaluate option D. Forwarding requires additional hardware: multiplexers at the ALU inputs to select between register file data and forwarded data, plus dedicated wiring (forwarding paths) from pipeline registers to those multiplexers. It is not free. Option D is INCORRECT.

    Answer: A, B, C

    Question 10 · Computer Organization and Architecture · 2023 NAT
    Consider a 3-stage pipelined processor having a delay of 10 ns (nanoseconds), 20 ns, and 14 ns, for the first, second, and the third stages, respectively. Assume that there is no other delay and the processor does not suffer from any pipeline hazards. Also assume that one instruction is fetched every cycle.

    The total execution time for executing 100 instructions on this processor is __________ ns.
    Correct Answer:

    2040

    Step-by-Step Solution

    Insight: In a synchronous pipeline, the clock cycle time is dictated by the slowest stage. Total time is (k + n - 1) multiplied by this cycle time.

    Exam route: Max stage delay = 20 ns. Latch delay = 0. Cycle time = 20 ns. Total cycles = 3 + 100 - 1 = 102. Total time = 102 × 20 = 2040 ns.

    Learning route:

    1. Identify the delay of each stage: 10 ns, 20 ns, 14 ns.
    2. Determine the pipeline cycle time, which is the maximum stage delay plus any latch delay. Here, max(10, 20, 14) + 0 = 20 ns.
    3. Identify the number of stages (k = 3) and the number of instructions (n = 100).
    4. Calculate the total number of cycles required to execute n instructions in a k-stage pipeline: k + n - 1 = 3 + 100 - 1 = 102 cycles.
    5. Multiply the total cycles by the cycle time: 102 × 20 ns = 2040 ns.

    More previous year questions (pyqs) in this unit