luyanaa's recent timeline updates
luyanaa's repos on GitHub
C++ · 5 watchers
armor2022
A RoboMaster 2022 revised version of auto-aiming codes from Tongji University Team SuperPower.
MATLAB · 3 watchers
crispy-enigma
Repo to store my MIT 9.40 'Introduction to Neural Computation' notes and assignments.
C · 3 watchers
openwrt-duo
OpenWRT for CVITEK CV1800B and VexRISCV LiteX
SourcePawn · 3 watchers
TR-1um
An experiment to port TR-1um and OpenSUSI MPW template to LibreLane + ALIGN.
Python · 2 watchers
flash-attn-triton
Flash Attention v2 (and SageBwd) Implementation targeted for NVIDIA Volta and Turing.
Verilog · 2 watchers
icsprout55-pdk
Clean-Room Analog Reverse-Engineering on ICsprout 55nm, refer to PTM BSIM4 and FreePDK45.
Python · 2 watchers
siliconcraft
siliconcraft — an open-EDA PDK from the NCSU CDK
Python · 1 watchers
NeoBERT
A reference BERT with modern techniques, forked from NeoBERT.
C++ · 1 watchers
neoSYCL
A SYCL Implementation for CPU and SX-Aurora TSUBASA (fork for newer LLVM, WIP)
Python · 1 watchers
rebuilded_classifier
1 watchers
resume
C++ · 0 watchers
2021RM-OPENCV
Python · 0 watchers
ALIGN-public
Analog Layout, Intelligently Generated from Netlists for TR-1um
C · 0 watchers
bf89
bf89 - A Simple BrainF**k Interpreter In Pure C89
MATLAB · 0 watchers
bookish-waffle
Handout of 12211402 Mathematical Experiment
Python · 0 watchers
CausalEEG
MATLAB · 0 watchers
gKDR-GMM
gradient kernel dimension reduction combined with gaussian mixture model for neural network modeling
Python · 0 watchers
glowing-journey
The Bayesian C. elegans Project
Lua · 0 watchers
kilo-code.nvim
A Neovim plugin that integrates KiloCode CLI into Neovim, providing a sidebar interface and automatic file change detection.
MATLAB · 0 watchers
laughing-train
Python · 0 watchers
legendary-octo-parakeet
EEG experiments
C++ · 0 watchers
LightGBM
Fix LightGBM int64_t index based on https://github.com/microsoft/LightGBM/pull/4809/files, Pass Unit Test.
Pascal · 0 watchers
lottery-newyear2017
HTML · 0 watchers
luyanaa
Ruby · 0 watchers
luyanaa.github.io
Python · 0 watchers
magi
cross-species brain foundation model.
Python · 0 watchers
mistofrutta
Collection of random utilites
0 watchers
nix-eda-env
Python · 0 watchers
OpenRAM
An open-source static random access memory (SRAM) compiler for icsprout55.
Python · 0 watchers
pumpprobe
Signal propagation atlas libraries, Randi et al. 2022 (fork by Robin Lu)
C++ · 0 watchers
RaceBot
Tongji University CS100433: Computer Graphics (2021 Fall)
0 watchers
RhyMAX
It is a systolic array, forked from https://github.com/LeonZ03/AME. It was originally developed by RACE and now is used in XSAI
Shell · 0 watchers
rocm-port
Forked from rocm-loongarch, ROCm build scripts for alternative architectures including LoongArch and RISC-V.
Python · 0 watchers
savitzkygolay
Savitzky-Golay filters
Python · 0 watchers
studious-tribble
Handout of Numerical Computation (10044202)
Scala · 0 watchers
SynapseML
Simple and Distributed Machine Learning
0 watchers
tarsier-monthly
Monthly Progress Reports for RISC-V Operating Systems
Python · 0 watchers
TDE-RICA
Refactor and Enhancement to original TDE-RICA.
PHP · 0 watchers
tinyquant
Tiny quant bot for PHP web-hosting service.
Python · 0 watchers
TR-1um_DRC_Regression_TEST
KLayout DRC runset regression project
Python · 0 watchers
TR-1um_MPW_template
C++ · 0 watchers
triton
Triton 3.2 for Volta and Turing.
MATLAB · 0 watchers
upgraded-octo-enigma
Repo to store homework of my mathematical modeling course.
Python · 0 watchers
wormbrain
Handling of small brains (fork by Robin Lu)
Python · 0 watchers
wormdatamodel
Data model for whole brain recordings (fork by Robin Lu)
luyanaa

luyanaa

V2EX member #191913, joined on 2016-09-15 23:52:32 +08:00
luyanaa's recent replies
Jun 2, 2020
Replied to a topic by Eender › Python › AMD 跑深度学习
印象里面 Github 上面 ROCm 跑 ResNet50 的 benchmark 相对价格不算很难看,主要是受 AMD 每代产品定位的拖累( Radeon VII 毕竟买的人实在太少,大多数人买的 Vega56/Vega64(GFX9 架构)或者 RX580(GFX8)说到底就是甜品卡,就算是平常的使用环境 RX580 也只能和 GTX1050Ti 或者 GTX1060 对比,Vega64 只能和 1070Ti 对比)。实际上我个人的统计(仅供参考,不保证完全控制变量,不保证实际使用体验,来源基本都是 Github 的 issue 和 lambdalabs 的测试数据) Resnet50 benchmark Vega64 比 1080Ti 慢 1/7,Radeon VII 甚至 ROCm2.7 能够接近 RTX2080Ti 的表现。但显然 ROCm 各方面的支持做的不好,新架构的支持偏慢( RDNA 我印象里面似乎还只有 unofficial 的 port,只是填了基础的坑,离开箱可用还差一些),性能还得一点点鸡血上去( Radeon VII 从 2.1.96 到 2.6 似乎 Image/sec 涨了快四分之一),等到满血了很可能下一代甚至下两代都已经出来了,而且动不动还有各种神奇的锅。(当然我以上的数据都只算了 ResNet50 的 Benchmark,因为这个 Benchmark 那个 issue 里测的最多,最方便进行有意义的统计,并不全面,但应该能反映一些问题)
知乎里面可能比较值得参考的几个帖子: https://www.zhihu.com/question/53091802/answer/890213654
https://zhuanlan.zhihu.com/p/80531243
听音乐,吃话梅(
Feb 7, 2018
Replied to a topic by boboliu › 分享创造 › yinshiGo - 又一个有点想法的一言服务端
有个微小的提议,第一句话能不能不要那么苟,容易翻车
About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Privacy   ·   Solana   ·   2535 Online   Highest 6679   ·     Select Language
创意工作者们的社区
World is powered by solitude
VERSION: 3.9.8.5 · 19ms · UTC 14:08 · PVG 22:08 · LAX 07:08 · JFK 10:08
♥ Do have faith in what you're doing.