The document presents a method for hierarchical object detection in images using deep reinforcement learning. It explores the formulation of the detection process as a Markov decision process and tests different feature extraction strategies, concluding that the image-zooms model provides better results, albeit with less accuracy in the final bounding boxes. Suggestions for improvement include training a regressor to fine-tune bounding box precision.