In the previous chapter, Chapter 3: AI Prompt Generation, we acted as the "Head Chef." We wrote a recipe (the Prompt) telling the AI what to do: "Fetch the comments."
But there is a problem. The AI is smart, but it doesn't inherently know where your files are or how to access GitHub's private servers. It needs a tool to reach the outside world.
In this chapter, we will define the GitHub Data Retrieval Strategy. We will give the AI a "Remote Control" (the GitHub CLI) and a specific set of buttons to press to get the data we need.
Imagine you go to a massive library to find a specific quote in a specific book.
If you just tell the librarian: "Get me the quote," they won't know where to look. You need to provide a Retrieval Strategy:
In our project:
gh commands we put in our prompt.Without this strategy, the AI acts blindly. With it, the AI becomes a surgical tool that extracts exactly the data we need.
To execute this strategy, we use two main concepts:
gh)
We don't ask the AI to write complex HTTP network requests. Instead, we ask it to use the GitHub CLI (gh).
This is a command-line tool installed on your computer. It handles authentication and connection details automatically. It is our "Remote Control" for GitHub.
GitHub stores comments in two different "buckets," and we need to look in both:
var to const on line 10").
We implement this strategy by writing specific instructions inside our AI Prompt (in index.ts). Let's break down the three steps we command the AI to take.
First, the AI needs to know which Pull Request (PR) it is currently looking at. It needs the PR number and the repository owner.
The Instruction:
1. Use `gh pr view --json number,headRepository`...
What it does: This command asks GitHub: "Tell me about the PR I am currently in."
number (e.g., #42) and the repository name.Now that the AI knows the PR number (let's say #42), it can fetch the general discussion.
The Instruction:
2. Use `gh api /repos/{owner}/{repo}/issues/{number}/comments`...
Why /issues/?
In GitHub's internal database, every Pull Request is technically also an "Issue." The general conversation lives in the "Issues" bucket.
Finally, we need the technical comments attached to specific lines of code.
The Instruction:
3. Use `gh api /repos/{owner}/{repo}/pulls/{number}/comments`...
Why /pulls/?
Code-specific comments (reviews) only exist on Pull Requests, not regular Issues. So we switch to the "Pulls" bucket to find them.
How does the AI actually execute this? It's a conversation between the AI and your Terminal.
Here is a simplified view of the AI gathering the ingredients:
Let's look at index.ts again to see exactly how these commands are embedded in the text string we return.
The Prompt Construction:
// index.ts
async getPromptWhileMarketplaceIsPrivate(args) {
return [
{
type: 'text',
text: `...
Follow these steps:
1. Use \`gh pr view --json number,headRepository\` to get the PR number...
Explanation:
We literally type the command inside backticks (\). The AI recognizes this pattern and knows it is a terminal command it should execute.
**The Advanced Fetch (Optional):**
Sometimes a comment says "Change this line," but we can't see the line. We give the AI a strategy for that too:
%%CODE_4%%
**Explanation:**
* gh api .../contents: This tells GitHub to send the actual file content.
* | jq ...: This pipes the result into a tool called jq to clean up the data.
* This instruction is conditional ("If the comment references..."). The AI decides when to use it.
## Conclusion
In this chapter, we defined the **GitHub Data Retrieval Strategy**.
We learned that:
1. We use the **gh CLI** as our remote control.
2. We fetch **Metadata** first to get the PR Number.
3. We fetch **General Comments** from the /issues/ endpoint.
4. We fetch **Review Comments** from the /pulls/` endpoint.
Now the AI has all the raw data it needs! But raw data is messy. It's just a giant pile of JSON text. We need to tell the AI how to present this information beautifully to the user.
Next Chapter: Output Formatting Specification
Generated by Code IQ