Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
hazrmard
82 days ago
|
parent
|
context
|
favorite
| on:
Is One Layer Enough? A Single Transformer Layer Ma...
Good work! I wonder if meta-learning can play a better role here compared to heuristics or hindsight. MAML requires hessians, but first-order MAML or Reptile variants could help apply layer-wise adjustments to learning rates.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: