Conversation
|
Hey @kpuru88, can you provide the official site for Shipwright? And the Github org under which these PRs were cloned for Shipwright's CRB run? |
|
Hi @ashleyzhang01 , |
@kpuru88 can you make the repos in that org public? |
|
Hi @ashleyzhang01, Would it be okay if I added you as a member of the organization? We’d like to keep this private for now. |
111c0ca to
a20652e
Compare
|
Hi @ashleyzhang01 , |
|
@kpuru88 yes, you can add me to the org |
|
@ashleyzhang01 : I've added you. |
|
@ashleyzhang01 : Please let me know if you haven’t received access yet. |
|
@kpuru88 i just joined the org but i can't see the repos/PRs that were used for the runs for the benchmark. to add you guys to the benchmark, i would need to be able to validate those, run your review tool on the benchmark myself, and verify that you have enough online usage. |
|
Thanks. |
|
@kpuru88 at least 600 public PRs across different repos that our online benchmark would be able to pick up |
|
@ashleyzhang01 : |
|
@kpuru88 it should be real customer usage from different organizations |
|
Hi @ashleyzhang01 : Regarding the 600 PRs, could we use Shipwright to review open-source ones? |
|
Hi @ashleyzhang01 , |
|
@kpuru88 I checked what our pipeline picks up from Shipwright, and this is what we see from shipwright-agent[bot]:
Let me know if this is incorrect or I'm missing anything. Previously I said:
But this seems like you guys are just running it yourself, rather than reflecting real usage, meaning we can't actually score and assess. Also, none of these are merged, and if there isn't human developer activity acting on the review comments, along with a merge, we can't analyze and score the reviews. |
Martian submission requirements
Current generated results
These results are not presented as a completed leaderboard submission until the two unchecked Martian requirements are satisfied.
Validation
uv run --frozen pytest— 29 passeduv run --frozen ruff check .— passed