Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,7 @@ For non-interactive installs, global installs, agent-specific installs, updates,
| **Holoscan Sensor Bridge** | Agent-ready skills for Holoscan Sensor Bridge devkit workflows, including demo environment bring-up, FPGA flashing for Lattice and VB1940 hardware, example application execution, QA test-plan automation, and support for configuring and using the Holoscan Sensor Bridge FPGA intellectual property (IP) core. | [`hsb-setup`](skills/hsb-setup), [`hsb-flash`](skills/hsb-flash), [`hsb-app`](skills/hsb-app), [`hsb-test`](skills/hsb-test), [`hsb-ip-def`](skills/hsb-ip-def), [`hsb-ip-packetizer`](skills/hsb-ip-packetizer), [`hsb-ip-create-top`](skills/hsb-ip-create-top) |
| **Isaac for Healthcare Workflows** | Agent-ready skills for Isaac for Healthcare agentic and catheter-navigation workflows, covering task authoring, data pipelines, policy training and validation, CT-derived digital twins, DRR rendering, and interactive catheter simulation. | [`i4h-workflow`](skills/i4h-workflow), [`i4h-workflow-setup`](skills/i4h-workflow-setup), [`i4h-workflow-create`](skills/i4h-workflow-create), [`i4h-workflow-scene-edit`](skills/i4h-workflow-scene-edit), [`i4h-workflow-dataset-teleop`](skills/i4h-workflow-dataset-teleop), [`i4h-workflow-dataset-replay`](skills/i4h-workflow-dataset-replay), [`i4h-workflow-dataset-mimic`](skills/i4h-workflow-dataset-mimic), [`i4h-workflow-dataset-annotate`](skills/i4h-workflow-dataset-annotate), [`i4h-workflow-dataset-convert`](skills/i4h-workflow-dataset-convert), [`i4h-workflow-finetune`](skills/i4h-workflow-finetune), [`i4h-workflow-validate`](skills/i4h-workflow-validate), [`i4h-workflow-e2e`](skills/i4h-workflow-e2e), [`i4h-lerobot-viz`](skills/i4h-lerobot-viz), [`i4h-catheter-navigation`](skills/i4h-catheter-navigation), [`i4h-catheter-navigation-setup`](skills/i4h-catheter-navigation-setup), [`i4h-catheter-navigation-digital-twin`](skills/i4h-catheter-navigation-digital-twin), [`i4h-catheter-navigation-render-drr`](skills/i4h-catheter-navigation-render-drr), [`i4h-catheter-navigation-viewport`](skills/i4h-catheter-navigation-viewport), [`i4h-catheter-navigation-smoke`](skills/i4h-catheter-navigation-smoke), [`i4h-catheter-navigation-e2e`](skills/i4h-catheter-navigation-e2e) |
| **Jetson BSP** | Agentic skills for setting up and customizing an NVIDIA Jetson Linux Board Support Package (BSP) — pick a target, prepare image and sources, customize IO (camera, PCIe, USB, pinmux, clocks, and more), then promote, flash, and validate. | [`jetson-build-source`](skills/jetson-build-source), [`jetson-customize-camera`](skills/jetson-customize-camera), [`jetson-customize-clocks`](skills/jetson-customize-clocks), [`jetson-customize-fan`](skills/jetson-customize-fan), [`jetson-customize-mgbe`](skills/jetson-customize-mgbe), [`jetson-customize-nvpmodel`](skills/jetson-customize-nvpmodel), [`jetson-customize-pcie`](skills/jetson-customize-pcie), [`jetson-customize-pinmux`](skills/jetson-customize-pinmux), [`jetson-customize-uphy`](skills/jetson-customize-uphy), [`jetson-customize-usb`](skills/jetson-customize-usb), [`jetson-derive-carrier`](skills/jetson-derive-carrier), [`jetson-download-bsp`](skills/jetson-download-bsp), [`jetson-flash-image`](skills/jetson-flash-image), [`jetson-generate-kb`](skills/jetson-generate-kb), [`jetson-init-image`](skills/jetson-init-image), [`jetson-init-source`](skills/jetson-init-source), [`jetson-init-target`](skills/jetson-init-target), [`jetson-link-docs`](skills/jetson-link-docs), [`jetson-optimize-memory`](skills/jetson-optimize-memory), [`jetson-print-bsp-info`](skills/jetson-print-bsp-info), [`jetson-promote-image`](skills/jetson-promote-image), [`jetson-quick-start`](skills/jetson-quick-start), [`jetson-set-target`](skills/jetson-set-target), [`jetson-validate-image`](skills/jetson-validate-image) |
| **Jetson Device** | Device-side agent skills for working with a live NVIDIA Jetson after boot — diagnostics, memory auditing, headless setup, inference memory tuning, LLM serving and benchmarking, packaging guidance, and speculative decoding. | [`jetson-diagnostic`](skills/jetson-diagnostic), [`jetson-headless-mode`](skills/jetson-headless-mode), [`jetson-inference-mem-tune`](skills/jetson-inference-mem-tune), [`jetson-llm-benchmark`](skills/jetson-llm-benchmark), [`jetson-llm-serve`](skills/jetson-llm-serve), [`jetson-memory-audit`](skills/jetson-memory-audit), [`jetson-package`](skills/jetson-package), [`jetson-print-device-info`](skills/jetson-print-device-info), [`jetson-speculative-decoding`](skills/jetson-speculative-decoding) |
| **Jetson Device** | Device-side agent skills for working with a live NVIDIA Jetson after boot — diagnostics, memory auditing, headless setup, inference memory tuning, LLM serving and benchmarking, packaging guidance, and speculative decoding. | [`jetson-diagnostic`](skills/jetson-diagnostic), [`jetson-headless-mode`](skills/jetson-headless-mode), [`jetson-inference-mem-tune`](skills/jetson-inference-mem-tune), [`jetson-llm-benchmark`](skills/jetson-llm-benchmark), [`jetson-llm-serve`](skills/jetson-llm-serve), [`jetson-memory-audit`](skills/jetson-memory-audit), [`jetson-package`](skills/jetson-package), [`jetson-print-device-info`](skills/jetson-print-device-info), [`jetson-speculative-decoding`](skills/jetson-speculative-decoding), [`jetson-video-benchmark`](skills/jetson-video-benchmark), [`jetson-video-capability`](skills/jetson-video-capability), [`jetson-video-pipeline`](skills/jetson-video-pipeline), [`jetson-video-recipe`](skills/jetson-video-recipe), [`jetson-video-setup`](skills/jetson-video-setup) |
| **Medical AI Skills** | Agent-ready medical AI skills built on MONAI for DICOM handling, NVIDIA-hosted medical imaging model workflows, segmentation, synthesis, and evidence-oriented evaluation. | [`dicom-metadata-extract`](skills/dicom-metadata-extract), [`dicom-series-preflight`](skills/dicom-series-preflight), [`dicom-series-to-volume`](skills/dicom-series-to-volume), [`nv-generate-ct-rflow`](skills/nv-generate-ct-rflow), [`nv-generate-mr`](skills/nv-generate-mr), [`nv-generate-mr-brain`](skills/nv-generate-mr-brain), [`nv-generate-mr-brain-finetune`](skills/nv-generate-mr-brain-finetune), [`nv-generate-vae-finetune`](skills/nv-generate-vae-finetune), [`nv-reason-cxr`](skills/nv-reason-cxr), [`nv-segment-ct`](skills/nv-segment-ct), [`nv-segment-ct-finetune`](skills/nv-segment-ct-finetune), [`nv-segment-ctmr`](skills/nv-segment-ctmr) |
| **Megatron-Core** | Large-scale distributed training — model parallelism, pipeline parallelism, and mixed precision. | [`mcore-create-issue`](skills/mcore-create-issue), [`mcore-linting-and-formatting`](skills/mcore-linting-and-formatting), [`mcore-run-on-slurm`](skills/mcore-run-on-slurm), [`mcore-split-pr`](skills/mcore-split-pr), [`mcore-testing`](skills/mcore-testing) |
| **NeMo AutoModel** | NeMo AutoModel - PyTorch-native distributed training for LLMs/VLMs with Hugging Face support, recipes, launchers, and validation workflows. | [`nemo-automodel-distributed-training`](skills/nemo-automodel-distributed-training), [`nemo-automodel-launcher-config`](skills/nemo-automodel-launcher-config), [`nemo-automodel-model-onboarding`](skills/nemo-automodel-model-onboarding), [`nemo-automodel-recipe-development`](skills/nemo-automodel-recipe-development) |
Expand Down
98 changes: 98 additions & 0 deletions skills/jetson-video-benchmark/BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
# Skill Benchmark: jetson-video-benchmark

> ✅ **Overall verdict: PASS — Recommended for publication**

## Publication Recommendation

Recommended for publication based on the completed evaluation evidence in this report.

## Evaluation Metadata

- Skill: `jetson-video-benchmark`
- Evaluation date: 2026-08-10
- Evaluator version: `1.1.2`
- Agents: Claude Code (`aws/anthropic/bedrock-claude-opus-4-8`), Codex (`openai/openai/gpt-5.5`)
- Tasks: 7 evaluation tasks (7 positive)
- Dataset digest: `sha256:e2d80ff3bff2819a83593a6167c1fd9119caa0d8febe727d4a947e1b84baec88` (skill-evaluator-dataset-snapshot/1)
- Attempts per task: 1
- Environment: `k8s-sandbox`
- Tier 3 evidence: required for publication

Each task attempt ran in its own isolated sandbox pod.

## What This Report Answers

The three-tier evaluation checks whether the skill:

- is safe to use;
- produces correct answers;
- is discovered and activated when needed;
- helps the agent complete the user's goal and expected workflow; and
- avoids wasted skill and tool usage.

## Results at a Glance

| Measure | Claude Code (Baseline → Skill Uplift) | Codex (Baseline → Skill Uplift) |
|---|---:|---:|
| Overall | 47% → 97% (+50 points) | 49% → 92% (+43 points) |
| Security | 100% → 100% (±0 points) | 86% → 100% (+14 points) |
| Correctness | 51% → 100% (+49 points) | 69% → 83% (+14 points) |
| Discoverability | 28% → 100% (+72 points) | 48% → 94% (+46 points) |
| Effectiveness | 36% → 84% (+48 points) | 32% → 90% (+58 points) |
| Efficiency | 19% → 100% (+81 points) | 11% → 92% (+82 points) |

**How to read this table:** baseline is the same task attempted without the target skill. Uplift is `skill score - baseline score`, shown in percentage points.

Example: `47% → 92% (+45 points)` means the skill-assisted run scored 92%, 45 percentage points above its 47% no-skill baseline.

## Tier Status

| Tier | Purpose | Status | Evidence |
|---|---|---|---|
| Tier 1 | Static validation | **PASSED WITH OBSERVATIONS** | 1 validator(s); 1 finding(s) |
| Tier 2 | Semantic deduplication | **NOT RUN** | No result was recorded |
| Tier 3 | Live agent evaluation | **PASS** | 2 agent(s); 7 task(s) |

## Findings and Observations

<details>
<summary>Show detailed findings and successful checks</summary>

- **MEDIUM** SCHEMA/body_recommended_section: Missing recommended section: '## Examples' (`skills/jetson-video-benchmark/SKILL.md`)

</details>

## Scoring Methodology

<details>
<summary>Show dimension definitions, source signals, and thresholds</summary>

| Dimension | Question | Scored signals |
|---|---|---|
| Security | Is it safe to use? | `security` (100%) |
| Correctness | Is the answer correct? | `accuracy` (100%) |
| Discoverability | Was the right skill loaded when needed? | `skill_execution` (100%) |
| Effectiveness | Did the skill help complete the task? | `goal_accuracy` (50%) + `behavior_check` (50%) |
| Efficiency | Did it avoid wasted tool or skill usage? | `skill_efficiency` (100%) |

- Dimension bands: PASS at 50% or above; NEUTRAL from 40% to below 50%; FAIL below 40%.
- Overall Tier 3 lift: PASS at +5 points or more; FAIL at -10 points or less; values between those bands are NEUTRAL.
- Overall verdict: PASS only when every configured dimension passes for at least one supported agent. Lift is reported as diagnostic evidence and does not override this gate.
- The 50% attempt pass threshold is a separate per-task gate; it is not the dimension pass threshold.
- Effectiveness is the equal-weight mean of goal completion (`goal_accuracy`) and expected workflow adherence (`behavior_check`).
- Token efficiency is a separate report-only signal. It does not change a dimension score or the overall verdict.

Signals present in this run:

- `security` (Security): unsafe operations, secret leakage, and unauthorized access.
- `skill_execution` (Skill Execution): whether the expected skill was found and executed.
- `skill_efficiency` (Efficiency): routing quality, workspace-aware skill reads, and productive tool use.
- `accuracy` (Accuracy): final-answer correctness against the reference answer.
- `goal_accuracy` (Goal Accuracy): whether the user's goal was achieved.
- `behavior_check` (Behavior Check): whether the expected workflow behavior was followed.

</details>

## Freshness

Regenerate this benchmark when the skill, evaluation dataset, target agent/model, evaluator version, environment, or scoring policy changes.
Loading